AMD GPU Operator on Kubernetes — DaemonSets, device plugin, multi-tenant patterns

How the AMD GPU Operator builds on the NVIDIA GPU Operator playbook, what's different (cert-manager hard-required, KMM-based driver, no MIG on MI300X), how to install via helm, validate, and operate at scale.

Try the commands on this page in the command emulator — type help for the full list, or solutions for copy-paste fix recipes.

If you have run the NVIDIA GPU Operator on RKE2 / EKS / GKE / OpenShift in production, the AMD GPU Operator will feel like the second-system effect that learned a few lessons. AMD shipped its operator in 2024, several years after NVIDIA's, and the architecture reflects what AMD watched everyone else trip over: cert-manager is a hard prerequisite (no half-installs), the kernel module is managed via the upstream KMM operator (not a custom DaemonSet), and the DeviceConfig CRD gives you the "one chart, many node configurations" pattern out of the box.

It is also younger software. ROCm 7 plus AMD GPU Operator 1.4+ are getting close to NVIDIA-Operator-grade reliability, but you should expect more rough edges than on the NVIDIA side. This page covers what's there, how to install, how to validate, and the multi-tenant patterns that work today.

What the operator deploys

ComponentNVIDIA equivalentAMD component / DaemonSetWhat it does
Driver / kernel modulenvidia-driver-daemonsetKMM-managed Module CR (amdgpu)Loads amdgpu on each node
Device pluginnvidia-device-plugin-daemonsetamd-gpu-device-pluginAdvertises amd.com/gpu to kubelet
Node labeller / GFDgpu-feature-discoveryamd-gpu-node-labellerLabels nodes with GPU model/family
Container runtime hooksnvidia-container-toolkit-daemonset(none — relies on standard --device mounts)Different model — see below
Metricsnvidia-dcgm-exporteramd-device-metrics-exporterPrometheus metrics
Validatornvidia-operator-validator(validation pod per DeviceConfig)Checks the stack works end-to-end
MIG managernvidia-mig-manager(none — MI300X has no MIG; CDNA 4 partition manager planned)n/a yet

The biggest architectural divergence is how the kernel module gets onto the node. NVIDIA's operator ships its own privileged DaemonSet that runs make modules_install inside a container. AMD reuses KMM (Kernel Module Management), a Red Hat–originated upstream Kubernetes operator that handles kernel module lifecycle generically. KMM is more flexible (you can sign modules, you can target specific kernel versions, you can blacklist conflicting in-tree modules) but has more moving parts.

The other divergence: there is no AMD container-toolkit equivalent. NVIDIA injects a runtime hook into containerd that automatically mounts /dev/nvidia* and the userspace libraries. AMD relies on the device plugin advertising /dev/kfd and /dev/dri/render* and you mount the userspace as a normal volume (or bake it into your image). Less magic, more explicit.

Prerequisites

RequirementVersionWhy
Kubernetesv1.29+DeviceConfig CRD uses v1alpha1, requires recent k8s
Helmv3.2+Chart format
cert-managerv1.13+Mandatory. Operator's webhook needs TLS
Node Feature Discovery (NFD)v0.15+Auto-installed unless you set --set node-feature-discovery.enabled=false
Kernel Module Management (KMM)v2.0+Auto-installed unless --set kmm.enabled=false
Container runtimecontainerd 1.7+ or CRI-OStandard k8s runtimes; no AMD-specific hooks needed
Underlying nodesUbuntu 22.04 / 24.04, RHEL 9, kernel 5.14+ / 6.xSee ROCm for full matrix

The cert-manager requirement is non-negotiable. Don't skip it; the operator's admission webhook silently fails closed without TLS.

# Install cert-manager first
helm repo add jetstack https://charts.jetstack.io --force-update
helm repo update
helm install cert-manager jetstack/cert-manager \
  --namespace cert-manager \
  --create-namespace \
  --version v1.15.1 \
  --set crds.enabled=true

# Wait for cert-manager to be ready
kubectl wait --for=condition=Ready --timeout=120s pod \
  -n cert-manager -l app.kubernetes.io/instance=cert-manager

Two driver modes (same trade-off as NVIDIA)

Like the NVIDIA operator, the AMD operator supports two driver modes:

Mode A: pre-installed driver (spec.driver.enable: false)

The host already has amdgpu-dkms installed and amdgpu is loaded. The operator skips driver management and only deploys device-plugin / labeller / metrics.

This is the right choice for:

  • Bare-metal fleets where you bake drivers into the OS image.
  • Air-gapped clusters where pulling a driver-container image is awkward.
  • Mixed-vendor clusters where you want clear ownership boundaries.

Mode B: operator-managed out-of-tree driver (spec.driver.enable: true)

The operator (via KMM) loads the kernel module from a container image at pod startup. The host kernel doesn't need amdgpu-dkms installed.

This is the right choice for:

  • Greenfield clusters where you want declarative driver management.
  • Heterogeneous kernel versions across nodes that you don't want to handle manually.
  • Frequent ROCm version churn where bumping the operator chart is cleaner than re-imaging.

For most established fleets I've seen, Mode A (pre-installed) wins. You get more control, fewer surprises during outages, and the OS-level drift management you already have (Ansible / Puppet / OS image rebuilds) handles the driver consistently.

Helm install: the canonical path

# Add the AMD operator chart repo
helm repo add rocm https://rocm.github.io/gpu-operator
helm repo update

# Install
helm install amd-gpu-operator rocm/gpu-operator-charts \
  --namespace kube-amd-gpu \
  --create-namespace \
  --version=v1.4.1

That installs the operator itself. It does not yet do anything to your nodes — the operator is just a control plane. You also need to apply a DeviceConfig Custom Resource to tell it which nodes to target.

DeviceConfig: the per-pool configuration

DeviceConfig is the AMD operator's central CRD. One per "pool of nodes that share configuration" — usually one for the whole cluster, but you can have multiple if you have heterogeneous nodes.

Example: pre-installed driver, default plugin, metrics on

apiVersion: amd.com/v1alpha1
kind: DeviceConfig
metadata:
  name: production-gpus
  namespace: kube-amd-gpu
spec:
  driver:
    enable: false                   # host has amdgpu-dkms already
  devicePlugin:
    devicePluginImage: rocm/k8s-device-plugin:latest
    nodeLabellerImage: rocm/k8s-device-plugin:labeller-latest
  metricsExporter:
    enable: true
    serviceType: "ClusterIP"        # or NodePort if scraping from outside
    image: docker.io/rocm/device-metrics-exporter:v1.4.1
    rbacConfig:
      enable: true                  # service account + RBAC for the exporter
  selector:
    feature.node.kubernetes.io/amd-gpu: "true"

Apply it:

kubectl apply -f deviceconfig.yaml

The selector says: target only nodes that NFD has labelled amd-gpu: "true". If NFD isn't installed (you set --set node-feature-discovery.enabled=false), use a manual label:

kubectl label node mynode amd.com/gpu-node=true

And update the selector accordingly.

Example: operator-managed driver

apiVersion: amd.com/v1alpha1
kind: DeviceConfig
metadata:
  name: greenfield-gpus
  namespace: kube-amd-gpu
spec:
  driver:
    enable: true
    blacklist: true                  # blacklist any in-tree amdgpu the host might load
    image: docker.io/rocm/amdgpu-driver:30.20.1-22.04
    version: "30.20.1"
  devicePlugin:
    devicePluginImage: rocm/k8s-device-plugin:latest
    nodeLabellerImage: rocm/k8s-device-plugin:labeller-latest
  metricsExporter:
    enable: true
    serviceType: "ClusterIP"
    image: docker.io/rocm/device-metrics-exporter:v1.4.1
  selector:
    feature.node.kubernetes.io/amd-gpu: "true"

The blacklist: true is important — without it, on a node that already has the in-tree amdgpu loaded by systemd-udev, the operator will fail to load its own version. Blacklist forces a clean slate.

Validation: end-to-end

Same flow as NVIDIA: install, validate, run a smoke pod.

# 1. Operator pods running?
$ kubectl get pods -n kube-amd-gpu
NAME                                                   READY   STATUS    RESTARTS   AGE
amd-gpu-operator-controller-manager-xxx                2/2     Running   0          2m
amd-gpu-operator-device-plugin-xxx                     1/1     Running   0          1m
amd-gpu-operator-node-labeller-xxx                     1/1     Running   0          1m
amd-gpu-operator-metrics-exporter-xxx                  1/1     Running   0          1m

# 2. Nodes have amd.com/gpu allocatable?
$ kubectl describe node my-mi300x-node | grep -i amd.com/gpu
  amd.com/gpu:  8

# 3. Or as a custom-columns sweep:
$ kubectl get nodes -o custom-columns=NAME:.metadata.name,GPU:'.status.allocatable.amd\.com/gpu'
NAME                       GPU
my-mi300x-node-0           8
my-mi300x-node-1           8
my-mi300x-node-2           8

# 4. DeviceConfig status
$ kubectl get deviceconfigs -n kube-amd-gpu -o yaml | grep -A 10 status:

# 5. Run a smoke pod
$ kubectl run amd-smi-test --rm -it --restart=Never \
  --image=docker.io/rocm/pytorch:latest \
  --overrides='{"spec":{"containers":[{"name":"amd-smi-test","image":"docker.io/rocm/pytorch:latest","command":["bash","-c","amd-smi monitor -ptu && rocminfo | grep gfx"],"resources":{"limits":{"amd.com/gpu":"1"}}}]}}'

If the smoke pod sees one GPU at the right gfx target, you are end-to-end functional.

The full smoke-pod manifest

apiVersion: v1
kind: Pod
metadata:
  name: amd-rccl-test
  namespace: default
spec:
  restartPolicy: Never
  containers:
  - name: amd-rccl-test
    image: docker.io/rocm/pytorch:latest
    command: ["/bin/bash", "-c"]
    args:
      - |
        set -e
        echo "=== rocm-smi ==="
        rocm-smi --showid
        echo "=== rocminfo gfx ==="
        rocminfo | grep -E "Name:.*gfx"
        echo "=== xGMI topo ==="
        rocm-smi --showtopo
        echo "=== PyTorch sanity ==="
        python -c "import torch; print('CUDA available (RCCL via NCCL backend):', torch.cuda.is_available()); print('Device count:', torch.cuda.device_count()); print('Arch list:', torch.cuda.get_arch_list())"
    resources:
      limits:
        amd.com/gpu: 8
      requests:
        amd.com/gpu: 8

For multi-node validation, run an MPI-launched rccl-tests job using Volcano or the Kubeflow MPIOperator — the same pattern that runs nccl-tests on NVIDIA fleets, just with the AMD container image and the amd.com/gpu resource type.

Resource model: what amd.com/gpu actually means

A pod requests amd.com/gpu: 1 and gets:

  • One GPU, exposed via /dev/dri/renderN plus /dev/kfd.
  • Read/write group permissions configured by the device plugin.
  • HSA-visible to the runtime; rocminfo and rocm-smi will show only the assigned GPU.

The Kubernetes scheduler treats it as opaque-extended-resource — same model as nvidia.com/gpu. No fractional resources, no priority among consumers within a node.

What it does not do (yet) on MI300X:

  • No fractional GPUs. You cannot request amd.com/gpu: 0.5. This is the same constraint NVIDIA had pre-MIG.
  • No GPU isolation between containers. Two containers requesting amd.com/gpu: 1 each on a 2-GPU node will get one each, but if one of them sets HIP_VISIBLE_DEVICES=0,1 it can see and use both. The device plugin restricts what the kernel exposes; it does not enforce HSA-level isolation.

Multi-tenant patterns: what works on MI300X today

The most-asked question coming from NVIDIA: "I'm used to slicing my H100 into 7 MIG instances for cheap inference. How do I do that on MI300X?"

Short answer: you don't, on MI300X. CDNA 3 has no MIG equivalent. The MI300X is a single GPU, atomically. You cannot slice it.

Slightly longer answer: CDNA 4 (MI350X / MI355X) has partitioning modes, but they are coarser than MIG and the operator support is younger. Specifically:

ModeWhat it doesEquivalent in NVIDIA terms
SPX (Single Partition)One GPU = one partition. The default.Default H100 mode
CPX (Compute Partition)Two compute partitions per OAM module.(rough) MIG 1g.40gb-style, 2 instances
TPX (Triple)Three compute partitions(rough) more aggressive slicing
QPX (Quad)Four compute partitions(rough) MIG 1g.10gb / 1g.20gb

These are runtime-configurable on CDNA 4 silicon (a rocm-smi call to the partition controller, then a node reboot). The AMD GPU Operator is gaining a "GPU Partitioning Manager" component analogous to NVIDIA's nvidia-mig-manager — at time of writing this is in early access.

For MI300X today, the practical multi-tenant patterns are:

Pattern 1: One container = one whole GPU

The simplest model. Workloads request amd.com/gpu: 1 and get a whole GPU. Inference services that don't need 192 GB of HBM still consume a whole GPU's worth of scheduler slot. This is fine for high-throughput, low-tenant-count clusters (one model per GPU, batch size large).

Pattern 2: Time-sliced via Triton / vLLM / TGI

If you're serving multiple inference models on one GPU, use a per-GPU multi-tenant inference server like vLLM with multi-LoRA, or NVIDIA Triton (yes, it works on AMD), or TGI (HuggingFace) with per-request routing. The Kubernetes scheduler still sees one amd.com/gpu, but the inference server multiplexes inside.

This is the pattern most production AMD inference deployments use today. It's how Microsoft Azure runs Llama models on MI300X SKUs at scale.

Pattern 3: One node, many containers, careful HIP_VISIBLE_DEVICES

If you have an 8-GPU node and want to run 8 different small inference services, pin each via HIP_VISIBLE_DEVICES=N and request amd.com/gpu: 1 per pod. The scheduler will place 8 pods on the node; each gets a different GPU. This works but requires the workload to behave when it sees only its assigned GPU.

Pattern 4: Wait for MI355X partitioning

For genuine MIG-style slicing — multiple isolated tenants on one GPU — wait for CDNA 4 partition support to mature in the operator. As of mid-2026, basic CPX support is shipping; TPX/QPX needs more bake time.

Metrics and observability

The AMD device-metrics-exporter is the analog of dcgm-exporter. Same Prometheus exposition format, similar label conventions:

# A typical metric line
amd_gpu_temperature_celsius{gpu="0",hostname="my-mi300x-node-0",model="MI300X"} 38.0
amd_gpu_power_watts{gpu="0",hostname="my-mi300x-node-0",model="MI300X"} 142.0
amd_gpu_memory_used_bytes{gpu="0",hostname="my-mi300x-node-0",model="MI300X"} 5.36e+08
amd_gpu_memory_total_bytes{gpu="0",hostname="my-mi300x-node-0",model="MI300X"} 2.06e+11
amd_gpu_xgmi_link_error_count{gpu="0",link="0",hostname="my-mi300x-node-0"} 0

To hook into Prometheus:

apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata:
  name: amd-device-metrics-exporter
  namespace: kube-amd-gpu
spec:
  selector:
    matchLabels:
      app: amd-device-metrics-exporter
  endpoints:
  - port: metrics
    interval: 15s

The metrics worth alerting on:

MetricAlert when
amd_gpu_uncorrectable_errors_total> 0 (any UE → drain node)
amd_gpu_temperature_celsius> 95 sustained (thermal limit approaching)
amd_gpu_throttle_statusTHERM or PROCHOT for > 60s
amd_gpu_xgmi_link_error_countrate-of-change > 0 (link degrading)
amd_gpu_memory_used_bytes / total_bytes> 0.95 (OOM risk for the workload)
amd_gpu_power_watts< 100 W during workload (GPU idle when expected busy)

Anyone running NVIDIA DCGM dashboards can port them with a sed-script — the metric names are different but the conceptual list is identical.

Operational patterns at scale

Node draining for maintenance

# Same as any k8s drain, but be aware that GPU pods may have to spin up large model weights
# from network storage on a different node — drain windows are longer than CPU-only

kubectl drain my-mi300x-node-3 --ignore-daemonsets --delete-emptydir-data

# When done:
kubectl uncordon my-mi300x-node-3

Driver upgrade via the operator

In Mode B (operator-managed driver), bumping the ROCm version is:

spec:
  driver:
    enable: true
    image: docker.io/rocm/amdgpu-driver:30.20.2-22.04
    version: "30.20.2"

The operator will roll the driver one node at a time (controlled by KMM's update strategy). Pods on a node draining for driver update will be evicted; if your training job spans the whole cluster, you need to coordinate the upgrade with the application.

In Mode A (pre-installed driver), the upgrade flow is the OS image / DKMS path — totally outside the operator. See ROCm.

Heterogeneous fleets (mixed AMD models)

If you have a mix of MI250X (gfx90a), MI300X (gfx942), and MI355X (gfx950) in the same cluster, you need:

  • The node-labeller component, which adds amd.com/gpu.product=MI300X style labels.
  • Workload selectors that match the right model:
spec:
  nodeSelector:
    amd.com/gpu.family: "Instinct"
    amd.com/gpu.product: "MI300X"
  • Container images compiled for the right gfx targets, OR a multi-arch image with PYTORCH_ROCM_ARCH=gfx90a;gfx942;gfx950.

Mixing AMD and NVIDIA in one cluster

Yes, you can do this. They're independent resources (nvidia.com/gpu vs amd.com/gpu), independent operators, independent device plugins. The risks:

  • Both operators install NFD and KMM. Coordinate --set node-feature-discovery.enabled=false and --set kmm.enabled=false on whichever you install second.
  • Some monitoring stacks assume NVIDIA-only metric names. Verify your Grafana dashboards work with both.
  • Workload schedulers (Volcano, Kueue, MCAD) need to be configured to understand both resource types.

This is increasingly common in mixed-vendor clouds. It works.

What's missing today (the honest list)

Things the AMD GPU Operator does not have, that the NVIDIA Operator does:

  • Mature MIG-equivalent. CDNA 3 has nothing; CDNA 4 partitioning is shipping but rough.
  • Confidential computing flow. NVIDIA has GPU TEE / Confidential VMs in production. AMD's equivalent (SEV-SNP + GPU TEE) is in earlier stages on the operator side.
  • GPU Direct Storage operator integration. NVIDIA bundles GDS configuration; AMD users do it manually.
  • Fabric-manager component. Not strictly needed (no NVSwitch), but the lack of a single-pane-of-glass for IF link health is felt — you query each node individually.
  • Time-slicing config in the operator. NVIDIA supports replicas: N to oversubscribe a GPU; AMD's analog is in the device plugin but not as well-integrated into the operator chart.

These will close over time. None are hard blockers for production today; they're papercuts.

See also