Skip to main content
Blog

BETA DETECTION: Inference Workload Pod Querying the Kubernetes API Analytic

  • October 7, 2026
  • 0 replies
  • 8 views
Aaron Beardslee
Forum|alt.badge.img
name: Inference Workload Pod Querying the Kubernetes API Analytic
signatureid: CLO-KUB01-RUN
category: 'Discovery'
threatname: 'Container and Resource Discovery'
functionality: 'Containers As A Service'

description: |
Detects a model-serving pod using its own service account token to query
the Kubernetes API.

An inference workload has no operational reason to talk to the control
plane. It loads a model, serves predictions, and reads nothing from the
cluster it runs in. When such a pod begins asking the API what it is
permitted to do, the most likely explanation is that something other than
the model server is now executing inside the container.

This matters for AI-as-a-Service infrastructure in particular. Model files
in the PyTorch and pickle formats are executable content: loading one runs
whatever code it carries. Wiz Research demonstrated the full consequence
against Hugging Face, uploading a modified gpt2 model that achieved remote
code execution inside the inference pod, then reaching the EC2 Instance
Metadata Service, minting a Kubernetes token with the node role, and
reading secrets belonging to other tenants.

The deserialization itself produces no control plane telemetry at all. No
Kubernetes API call is made, so no audit event exists, and the attack is
undetectable from audit logs by construction. What IS visible is the
consequence: once an attacker is executing in the container, orientation
follows, and orientation means asking the API questions.

The analytic therefore detects the attack by its effect rather than its
mechanism, and requires no runtime sensor to do so.

Two independent signals qualify the event. The Kubernetes audit record
names the originating pod, so the workload can be identified rather than
inferred. And the client identifies itself: legitimate in-cluster callers
use client-go and present a recognisable user agent, while a hand-rolled
request from inside a compromised container typically presents wget, curl
or a bare HTTP library.

reference:
- https://attack.mitre.org/techniques/T1613/
- https://attack.mitre.org/techniques/T1552/007/
- https://www.wiz.io/blog/wiz-and-hugging-face-address-risks-to-ai-infrastructure
- https://kubernetes.io/docs/reference/access-authn-authz/authentication/

labels:
- attack.discovery
- attack.t1613
- attack.credential_access
- attack.t1552.007

logsource:
product: AWS
service: AWS EKS Audit

detection:

selection_pod_identity:
AccountName|startswith: 'system:serviceaccount:'
# Raw:
# user.username
#
# Both AccountName and SourceUserName carry this. AccountName is used
# for consistency with the rest of the corpus; normalization uppercases
# it, which does not matter because Securonix policy comparisons are
# case insensitive.
#
# A pod acting with its mounted service account token, as opposed to a
# human or a node identity.

selection_inference_workload:
ObjectName|contains:
- 'model'
- 'inference'
- 'serving'
- 'predict'
# Raw:
# user.extra["authentication.kubernetes.io/pod-name"] -> customstring9
#
# Confirmed in the parser tester. NOT ContainerName, which exists in the
# dictionary but is not populated from this source.
#
# TUNE THIS LIST PER ENVIRONMENT. It is the naming convention of the
# inference workloads, and it is the only part of this analytic that is
# site specific. A deployment that names its pods differently will not
# match, and a deployment that names unrelated pods "model-..." will
# produce noise.

selection_orientation:
SourceServiceName:
- 'selfsubjectaccessreviews'
- 'selfsubjectrulesreviews'
- 'selfsubjectreviews'
# Raw:
# objectRef.resource
#
# "What am I allowed to do?" - the first question asked from a shell in
# a container, and never asked by a model server.

selection_handrolled_client:
RequestClientApplication|contains:
- 'Wget'
- 'curl'
- 'python-requests'
- 'python-urllib'
# Raw:
# userAgent
#
# Legitimate in-cluster clients use client-go and identify themselves,
# for example kubectl/v1.36 or a named controller. A bare downloader is
# a request somebody composed by hand inside the container.

condition: >
selection_pod_identity
and selection_inference_workload
and (selection_orientation or selection_handrolled_client)

criticality: High
saveasthreat: false


violation_summary:
grouping_attribute: 'AccountName'
level2_attribute: 'ObjectName'
level2_metadata_attributes:
metadata_attributes:

TECHNICAL DETAILS


    DATA SOURCE

    Amazon EKS control plane audit logs, delivered to CloudWatch Logs and
    collected by the AWS EKS Audit resource type over the awscloudwatch
    connector.

    Both EKS control plane log types write to a single log group,
    /aws/eks/<cluster>/cluster, and are separated only by log stream prefix.
    The audit connector must filter on the prefix kube-apiserver-audit-, and
    the authenticator connector on authenticator-. With the prefix left
    blank a connector ingests both, and authenticator records arriving at a
    JSON parser are rejected as unparsed.

    The datasource timezone must be UTC. Kubernetes audit timestamps are
    RFC3339 with an explicit Z, so a datasource configured for local time
    applies the offset a second time and events land in the future.


    PARSER SELECTION

    Securonix publishes two parsers for this resource type and the
    functionality is encoded in the parser name:

      SCNX_AMAZON_AWSEKSAUDIT_CAAS_AWS_JSO_COMM   Containers As A Service
      SCNX_AMAZON_AWSEKSAUDIT_CSA_AWS_JSO         Cloud Services / Applications

    This analytic requires the CAAS parser. The CSA variant maps to a
    functionality that does not carry DeviceAction, SourceServiceName or any
    container field, so the analytic cannot match against it.


    CONFIRMED FIELD MAPPING

    Validated in the Securonix parser tester against the raw event below,
    rather than assumed from the dictionary:

      user.username                        ->  SourceUserName (sourceusername)
      user.extra[...pod-name]              ->  ObjectName (customstring9)
      user.extra[...pod-uid]               ->  ReplicaSet (customstring44)
      userAgent                            ->  RequestClientApplication
      objectRef.resource                   ->  SourceServiceName
      annotations[...decision]             ->  DeviceAction
      user.groups                          ->  UserGroup (customstring10)
      sourceIPs                            ->  SourceAddress
      verb                                 ->  Method (requestmethod)
      requestObject...namespace            ->  Namespace (devicecustomstring1)

    NORMALIZATION HAPPENS AFTER THE PARSER, AND THE TESTER DOES NOT SHOW IT

    The parser tester maps user.username to SourceUserName and shows nothing
    writing AccountName. An ingested event nonetheless carries both:

      sourceusername   system:serviceaccount:h4x-demo:default
      accountname      SYSTEM:SERVICEACCOUNT:H4X-DEMO:DEFAULT

    So AccountName IS available, derived downstream of the parser and
    uppercased on the way. Reading the tester alone would wrongly suggest
    content written against AccountName cannot match.

    Normalization uppercases AccountName while SourceUserName preserves the
    emitted value. Securonix policy comparisons are case insensitive
    (confirmed by Aaron), so either field works and this analytic uses
    AccountName for consistency with the rest of the corpus.

    ContainerName is genuinely absent. It exists in the Containers As A
    Service dictionary but no field of the EKS audit record populates it;
    the pod name is in ObjectName.

    OBSERVED EVENT

    Captured on an EKS cluster built to reproduce the Hugging Face attack
    path, from a pod named model-server using the default service account of
    its namespace:

      user.username
        system:serviceaccount:h4x-demo:default

      user.extra
        authentication.kubernetes.io/pod-name   model-server
        authentication.kubernetes.io/node-name  ip-10-60-11-82.us-west-2...
        authentication.kubernetes.io/pod-uid    6c6225d1-8a1d-4e26-b064-...

      userAgent
        Wget

      objectRef.resource
        selfsubjectaccessreviews

      annotations["authorization.k8s.io/decision"]
        allow

    The same pod then attempted to list secrets, list pods cluster wide, and
    read selfsubjectreviews. All three were refused, which is the expected
    shape: orientation succeeds because self review is permitted to every
    authenticated principal, and the escalation that follows it fails.


    WHY THE INITIAL ACCESS STEP IS ABSENT FROM THE TRACE

    The malicious model was not executed. Pickle deserialization makes no
    Kubernetes API call and performs no authentication, so it generates
    neither an audit event nor an authenticator event. Reproducing it would
    have added nothing to this evidence. Detecting it requires runtime
    visibility inside the container, which is a different data source and
    different content.


    TUNING

    selection_inference_workload is the only site specific part. It should
    list the naming convention of the environment's model serving workloads.
    Where inference pods cannot be identified by name, a label based
    exclusion list of pods that are expected to call the API is the better
    inversion, though it is not expressible in this data source alone.

    selection_handrolled_client is deliberately an OR rather than a
    requirement. An attacker who uses a Kubernetes client library presents a
    normal user agent, and the pod identity alone should still fire.


Policy building walkthrough can be found in this previous post:

 

https://connect.securonix.com/threat%2Dresearch%2Dintelligence%2D62/beta%2Ddetection%2Dtelnyx%2Dteampcp%2Dcredential%2Dexfiltration%2D241