name: Inference Workload Pod Querying the Kubernetes API Analytic
signatureid: CLO-KUB01-RUN
category: 'Discovery'
threatname: 'Container and Resource Discovery'
functionality: 'Containers As A Service'
description: |
Detects a model-serving pod using its own service account token to query
the Kubernetes API.
An inference workload has no operational reason to talk to the control
plane. It loads a model, serves predictions, and reads nothing from the
cluster it runs in. When such a pod begins asking the API what it is
permitted to do, the most likely explanation is that something other than
the model server is now executing inside the container.
This matters for AI-as-a-Service infrastructure in particular. Model files
in the PyTorch and pickle formats are executable content: loading one runs
whatever code it carries. Wiz Research demonstrated the full consequence
against Hugging Face, uploading a modified gpt2 model that achieved remote
code execution inside the inference pod, then reaching the EC2 Instance
Metadata Service, minting a Kubernetes token with the node role, and
reading secrets belonging to other tenants.
The deserialization itself produces no control plane telemetry at all. No
Kubernetes API call is made, so no audit event exists, and the attack is
undetectable from audit logs by construction. What IS visible is the
consequence: once an attacker is executing in the container, orientation
follows, and orientation means asking the API questions.
The analytic therefore detects the attack by its effect rather than its
mechanism, and requires no runtime sensor to do so.
Two independent signals qualify the event. The Kubernetes audit record
names the originating pod, so the workload can be identified rather than
inferred. And the client identifies itself: legitimate in-cluster callers
use client-go and present a recognisable user agent, while a hand-rolled
request from inside a compromised container typically presents wget, curl
or a bare HTTP library.
reference:
- https://attack.mitre.org/techniques/T1613/
- https://attack.mitre.org/techniques/T1552/007/
- https://www.wiz.io/blog/wiz-and-hugging-face-address-risks-to-ai-infrastructure
- https://kubernetes.io/docs/reference/access-authn-authz/authentication/
labels:
- attack.discovery
- attack.t1613
- attack.credential_access
- attack.t1552.007
logsource:
product: AWS
service: AWS EKS Audit
detection:
selection_pod_identity:
AccountName|startswith: 'system:serviceaccount:'
# Raw:
# user.username
#
# Both AccountName and SourceUserName carry this. AccountName is used
# for consistency with the rest of the corpus; normalization uppercases
# it, which does not matter because Securonix policy comparisons are
# case insensitive.
#
# A pod acting with its mounted service account token, as opposed to a
# human or a node identity.
selection_inference_workload:
ObjectName|contains:
- 'model'
- 'inference'
- 'serving'
- 'predict'
# Raw:
# user.extra["authentication.kubernetes.io/pod-name"] -> customstring9
#
# Confirmed in the parser tester. NOT ContainerName, which exists in the
# dictionary but is not populated from this source.
#
# TUNE THIS LIST PER ENVIRONMENT. It is the naming convention of the
# inference workloads, and it is the only part of this analytic that is
# site specific. A deployment that names its pods differently will not
# match, and a deployment that names unrelated pods "model-..." will
# produce noise.
selection_orientation:
SourceServiceName:
- 'selfsubjectaccessreviews'
- 'selfsubjectrulesreviews'
- 'selfsubjectreviews'
# Raw:
# objectRef.resource
#
# "What am I allowed to do?" - the first question asked from a shell in
# a container, and never asked by a model server.
selection_handrolled_client:
RequestClientApplication|contains:
- 'Wget'
- 'curl'
- 'python-requests'
- 'python-urllib'
# Raw:
# userAgent
#
# Legitimate in-cluster clients use client-go and identify themselves,
# for example kubectl/v1.36 or a named controller. A bare downloader is
# a request somebody composed by hand inside the container.
condition: >
selection_pod_identity
and selection_inference_workload
and (selection_orientation or selection_handrolled_client)
criticality: High
saveasthreat: false
violation_summary:
grouping_attribute: 'AccountName'
level2_attribute: 'ObjectName'
level2_metadata_attributes:
metadata_attributes:TECHNICAL DETAILS
DATA SOURCE
Amazon EKS control plane audit logs, delivered to CloudWatch Logs and
collected by the AWS EKS Audit resource type over the awscloudwatch
connector.
Both EKS control plane log types write to a single log group,
/aws/eks/<cluster>/cluster, and are separated only by log stream prefix.
The audit connector must filter on the prefix kube-apiserver-audit-, and
the authenticator connector on authenticator-. With the prefix left
blank a connector ingests both, and authenticator records arriving at a
JSON parser are rejected as unparsed.
The datasource timezone must be UTC. Kubernetes audit timestamps are
RFC3339 with an explicit Z, so a datasource configured for local time
applies the offset a second time and events land in the future.
PARSER SELECTION
Securonix publishes two parsers for this resource type and the
functionality is encoded in the parser name:
SCNX_AMAZON_AWSEKSAUDIT_CAAS_AWS_JSO_COMM Containers As A Service
SCNX_AMAZON_AWSEKSAUDIT_CSA_AWS_JSO Cloud Services / Applications
This analytic requires the CAAS parser. The CSA variant maps to a
functionality that does not carry DeviceAction, SourceServiceName or any
container field, so the analytic cannot match against it.
CONFIRMED FIELD MAPPING
Validated in the Securonix parser tester against the raw event below,
rather than assumed from the dictionary:
user.username -> SourceUserName (sourceusername)
user.extra[...pod-name] -> ObjectName (customstring9)
user.extra[...pod-uid] -> ReplicaSet (customstring44)
userAgent -> RequestClientApplication
objectRef.resource -> SourceServiceName
annotations[...decision] -> DeviceAction
user.groups -> UserGroup (customstring10)
sourceIPs -> SourceAddress
verb -> Method (requestmethod)
requestObject...namespace -> Namespace (devicecustomstring1)
NORMALIZATION HAPPENS AFTER THE PARSER, AND THE TESTER DOES NOT SHOW IT
The parser tester maps user.username to SourceUserName and shows nothing
writing AccountName. An ingested event nonetheless carries both:
sourceusername system:serviceaccount:h4x-demo:default
accountname SYSTEM:SERVICEACCOUNT:H4X-DEMO:DEFAULT
So AccountName IS available, derived downstream of the parser and
uppercased on the way. Reading the tester alone would wrongly suggest
content written against AccountName cannot match.
Normalization uppercases AccountName while SourceUserName preserves the
emitted value. Securonix policy comparisons are case insensitive
(confirmed by Aaron), so either field works and this analytic uses
AccountName for consistency with the rest of the corpus.
ContainerName is genuinely absent. It exists in the Containers As A
Service dictionary but no field of the EKS audit record populates it;
the pod name is in ObjectName.
OBSERVED EVENT
Captured on an EKS cluster built to reproduce the Hugging Face attack
path, from a pod named model-server using the default service account of
its namespace:
user.username
system:serviceaccount:h4x-demo:default
user.extra
authentication.kubernetes.io/pod-name model-server
authentication.kubernetes.io/node-name ip-10-60-11-82.us-west-2...
authentication.kubernetes.io/pod-uid 6c6225d1-8a1d-4e26-b064-...
userAgent
Wget
objectRef.resource
selfsubjectaccessreviews
annotations["authorization.k8s.io/decision"]
allow
The same pod then attempted to list secrets, list pods cluster wide, and
read selfsubjectreviews. All three were refused, which is the expected
shape: orientation succeeds because self review is permitted to every
authenticated principal, and the escalation that follows it fails.
WHY THE INITIAL ACCESS STEP IS ABSENT FROM THE TRACE
The malicious model was not executed. Pickle deserialization makes no
Kubernetes API call and performs no authentication, so it generates
neither an audit event nor an authenticator event. Reproducing it would
have added nothing to this evidence. Detecting it requires runtime
visibility inside the container, which is a different data source and
different content.
TUNING
selection_inference_workload is the only site specific part. It should
list the naming convention of the environment's model serving workloads.
Where inference pods cannot be identified by name, a label based
exclusion list of pods that are expected to call the API is the better
inversion, though it is not expressible in this data source alone.
selection_handrolled_client is deliberately an OR rather than a
requirement. An attacker who uses a Kubernetes client library presents a
normal user agent, and the pod identity alone should still fire.
Policy building walkthrough can be found in this previous post:
