K8s GPU pod failures: from pod stuck to pod evicted
Decoder ring for every state a GPU pod can be stuck in — Pending, ContainerCreating, CrashLoopBackOff, OOMKilled, Evicted. What to look at, what to fix, and how to use ephemeral debug containers when exec doesn't work.
help for the full list, or solutions for reference procedures. Simulated outputs do not validate your environment.GPU pods fail in interesting ways. A normal pod failure is "image not found" or "OOM". A GPU pod failure can be any of those, plus device-plugin trouble, CDI spec mismatch, runtime class missing, IB device-plugin not exposing rdma resources, MIG slice mis-requested, GPU operator init crashlooping due to chart misalignment — and the symptom is "stuck Pending" or "ContainerCreating for 10 minutes".
This page is the symptom → cause map. Walk by phase: where is the pod stuck, what produces that phase, what to look at next.
The diagnostic recipe (always start here)
# 1. What state is the pod in?
$ kubectl get pod <name> -n <ns>
# 2. What does the scheduler / kubelet / runtime think?
$ kubectl describe pod <name> -n <ns>
# Read the Events: section bottom-up. The most recent event tells you the current state.
# 3. What did the container say (if it ran at all)?
$ kubectl logs <name> -n <ns>
$ kubectl logs <name> -n <ns> --previous # if it just restarted
$ kubectl logs <name> -n <ns> -c <container> # for multi-container or init containers
# 4. What does the node look like?
$ kubectl describe node <node> | head -50
$ kubectl get events --field-selector involvedObject.name=<node> -A
# 5. Get a shell or debug container if the pod is running
$ kubectl exec -it <name> -n <ns> -- bash
$ kubectl debug -it <name> -n <ns> --image=nicolaka/netshoot --target=<container>
For every section below, that's the entry point. The rest is interpretation.
Phase: Pending
The scheduler can't place the pod. Read kubectl describe pod Events: and look for FailedScheduling:
Warning FailedScheduling 30s default-scheduler 0/12 nodes are available:
8 Insufficient nvidia.com/gpu,
4 node(s) didn't match Pod's node affinity/selector.
| Predicate failure | Cause | Fix |
|---|---|---|
Insufficient nvidia.com/gpu | All matching nodes are at GPU capacity | Wait, scale, or schedule on a different pool |
didn't match Pod's node affinity/selector | nodeSelector / affinity rules don't match any node | Check labels: kubectl get nodes --show-labels |
had untolerated taint {key}={val} | Node has a taint the pod doesn't tolerate | Add toleration, or remove taint, or schedule elsewhere |
node(s) had volume node affinity conflict | PVC is bound to a zone the pod can't reach | Different PV / different scheduling target |
node(s) didn't match pod topology spread constraints | Anti-affinity / spread rules can't be satisfied | Loosen the constraint or scale up |
GPU-specific Pending reasons:
Insufficient nvidia.com/gpu
Verify the cluster has free GPU capacity:
$ kubectl get nodes -o custom-columns=NAME:.metadata.name,GPU:'.status.allocatable.nvidia\.com/gpu'
NAME GPU
gpu-node-01 8
gpu-node-02 8
gpu-node-03 8
$ kubectl describe node gpu-node-01 | grep -A 6 "Allocated resources"
Allocated resources:
Resource Requests Limits
cpu 64 128
memory 512Gi 1Ti
nvidia.com/gpu 8 8 # all GPUs in use
If everything's allocated, the scheduler is correct — wait, kill an old job, or scale.
Reservation mismatch
If you use Reservation / ReservationBinding (multi-tenant), pods can be Pending because:
- The pod doesn't match the binding's
podSelector→ no nodeAffinity injected → pod tries to land on any node, finds none with the right taint tolerated. - The binding selects but no reserved nodes are Ready or have free GPUs.
$ kubectl get reservations
$ kubectl get reservationbindings -n <tenant-namespace>
$ kubectl describe pod <name> -n <ns> | grep -A 4 Affinity
See Reservations.
Taints not tolerated
$ kubectl describe node <node> | grep Taints
Taints: nvidia.com/gpu=true:NoSchedule
reserved=tenant-foo:NoSchedule
Pod must tolerate every NoSchedule taint to land. Typical GPU node taints:
spec:
tolerations:
- key: nvidia.com/gpu
operator: Exists
effect: NoSchedule
- key: reserved
operator: Equal
value: tenant-foo
effect: NoSchedule
Phase: ContainerCreating
The pod has been scheduled to a node, but the runtime can't get the container running. kubectl describe pod Events: is the source of truth.
| Event | Cause | Fix |
|---|---|---|
Failed to pull image: ImagePullBackOff | Image not found, or registry creds wrong | Check imagePullSecrets, image tag, registry URL |
Failed to pull image: ErrImagePull (timeout) | Network to registry slow / blocked | Check egress firewall, mirror registry |
MountVolume.SetUp failed for volume "X": ... | PVC not bound, secret missing, configMap missing | kubectl get pvc, kubectl get secret/configmap -n <ns> |
failed to setup network for sandbox | CNI failure (IPAM exhausted, CNI daemon down) | kubectl logs -n kube-system <cni-pod> |
RunContainerError: failed to create containerd task | Runtime class missing or CDI spec missing | See below |
Pod sandbox changed, it will be killed and re-created | Kubelet restarted, transient | Usually self-heals |
CDI spec missing / runtime class not found
NVIDIA's Container Device Interface (CDI) is how new GPU Operator versions inject GPUs into containers. The kubelet needs to find a CDI spec that describes the GPU:
# on the node
$ ls /var/run/cdi/
nvidia.com-gpu.yaml
$ cat /var/run/cdi/nvidia.com-gpu.yaml | head -20
cdiVersion: 0.5.0
kind: nvidia.com/gpu
devices:
- name: "0"
containerEdits:
deviceNodes:
- path: /dev/nvidia0
...
If /var/run/cdi/ is empty:
- nvidia-cdi-hook isn't running → device-plugin / GPU Operator misconfigured.
- containerd config doesn't have CDI enabled.
# check containerd config
$ grep -A 3 "enable_cdi" /etc/containerd/config.toml
[plugins."io.containerd.grpc.v1.cri".containerd]
enable_cdi = true
cdi_spec_dirs = ["/etc/cdi", "/var/run/cdi"]
For runtime class:
# pod spec
spec:
runtimeClassName: nvidia # must exist in cluster
$ kubectl get runtimeclass
NAME HANDLER AGE
nvidia nvidia 127d
If nvidia runtime class is missing → device-plugin or GPU Operator missed it. kubectl get all -n gpu-operator.
Volume mounts (PVC stuck Pending)
$ kubectl get pvc -n <ns>
NAME STATUS VOLUME CAPACITY ACCESS MODES STORAGECLASS
my-pvc Pending weka-csi-block
$ kubectl describe pvc my-pvc -n <ns>
Events:
Warning ProvisioningFailed ... weka-csi-block: failed to provision: ...
Storage class issue (CSI driver down, quota exhausted, backend unreachable). The pod waits for the PVC; fix the PVC.
Phase: init container CrashLoopBackOff
Many GPU stacks have init containers that:
- Load kernel modules (privileged).
- Validate driver presence.
- Configure host directories.
- Wait for a CRD or operator state.
When these crashloop, the main container never starts.
The classic case from production: GPU Operator chart misalignment with what ArgoCD applied. ArgoCD reverts a container's startup args (e.g., the nvidia-driver-daemonset init container expects --kernel-version=$(uname -r) but ArgoCD is enforcing a different value from the manifest), and the init pod fails:
$ kubectl logs <init-pod> -c <init-container>
error: kernel version mismatch — host is 6.8.0-45-generic, image was built for 6.8.0-43
The fix isn't to keep restarting — it's to align the chart's manifest with what's actually deployed, or ignoreDifferences in ArgoCD if the diff is benign. See ArgoCD.
Other init container failures:
| Symptom | Cause |
|---|---|
failed to load nvidia kernel module | Kernel version != driver image's expected version |
failed to mount /run/nvidia/driver | Hostpath conflict; another driver pod is using it |
RDMA module load fails | OFED kernel modules missing on host |
init: waiting for CRD nvidia.com/clusterpolicy | Chart hasn't installed the CRD yet, or it was deleted |
kubectl logs <pod> -c <init-container> is the one source of truth here.
Phase: main container CrashLoopBackOff
Container starts, then exits non-zero. Look at logs:
$ kubectl logs <pod> -n <ns> --previous
GPU-specific patterns:
"GPU not visible inside the pod"
# from inside the container
$ nvidia-smi
No devices were found
Or in PyTorch:
>>> torch.cuda.device_count()
0
Yet the host has GPUs. Causes:
- device-plugin not running on the node.
kubectl get pods -n gpu-operator -o wide | grep device-plugin. If absent, nonvidia.com/gpuresource. - CDI spec missing (see above).
- Runtime class !=
nvidiain the pod spec. - Wrong number requested:
nvidia.com/gpu: 0in resources allocates none.nvidia.com/gpu: 8requires the node to have 8 free. - MIG slice not requested correctly: on a MIG-enabled GPU, you request
nvidia.com/mig-1g.10gb: 1notnvidia.com/gpu: 1. See MIG. NVIDIA_VISIBLE_DEVICES=voidin the image — sometimes set in CI base images. Override in the pod spec:env: - name: NVIDIA_VISIBLE_DEVICES value: all
Verification on a working pod:
$ kubectl exec -it <pod> -- nvidia-smi
+-----------------------------------------------------------------------------------------+
| NVIDIA-SMI ... |
| GPU ... |
| 0 NVIDIA H100 80GB HBM3 ... |
CUDA version mismatch
CUDA driver version is insufficient for CUDA runtime version
Container has CUDA 12.8 toolkit, host driver is 535 (CUDA driver API 12.2). See driver-firmware-mismatch.
Application-level
Look at logs for application error. NCCL errors get their own page: NCCL multi-node failures.
Phase: OOMKilled
$ kubectl describe pod <pod>
Last State: Terminated
Reason: OOMKilled
Exit Code: 137
Three flavors:
Host memory OOM (most common)
The container's cgroup memory limit was hit. kubectl get pod -o yaml and look at resources.limits.memory. Either bump the limit, or reduce the workload's memory footprint.
GPU workloads often pin host memory for staging buffers. A multi-node training job can pin several GB per rank. If your limits.memory is 256 GiB and the workload pins 280 GiB, it gets killed.
Pinned memory hitting cgroup limit
A subtle case: the workload mlocks buffers (RDMA needs this). The cgroup memory limit applies to mlocked memory too. If ulimits.l = unlimited but cgroup limit is 64 GiB, you can still hit 64 GiB and get OOMKilled even though your RSS appears normal.
Fix: raise resources.limits.memory to comfortably exceed expected pinning (NCCL buffer + framework staging + slack). For 8-GPU training, 1 TiB limit is a reasonable starting point.
GPU memory OOM (different — not OOMKilled)
A GPU OOM doesn't kill the container; it surfaces as a CUDA error inside the workload:
RuntimeError: CUDA out of memory. Tried to allocate 24.00 GiB ...
Reduce model batch size, use gradient checkpointing, or use torch.cuda.empty_cache() between phases.
GPU memory leaks: if nvidia-smi shows 80 GB used but no process is using it, you have a leaked CUDA context. nvidia-smi pmon to find the orphan, then kill it.
Phase: Evicted
Eviction is the kubelet protecting the node from going down. It's almost always one of:
$ kubectl describe node <node> | grep -A 8 Conditions:
Conditions:
MemoryPressure True ... kubelet has insufficient memory available
DiskPressure False
PIDPressure False
Ready True
| Condition True | Cause | Fix |
|---|---|---|
MemoryPressure | Node memory below threshold (default 100Mi free) | Restart leaky workloads, increase node memory, raise eviction threshold |
DiskPressure | Root or imagefs disk usage above threshold (default 85 %) | Clean up old images: crictl rmi --prune; clean /var/log |
PIDPressure | Too many processes | Find runaway pod with thousands of threads |
NetworkUnavailable | CNI dead | Restart CNI daemon |
Common GPU-host disk-pressure case: container images pile up in /var/lib/containerd/. Each driver image is several GB; after dozens of operator upgrades, you can fill the disk.
# clean unused images on a node (run via kubectl debug or ssh)
$ crictl images
$ crictl rmi --prune
# or for containerd's content store:
$ ctr -n k8s.io images ls
Also check /var/lib/kubelet/pods/ for leftover pod directories from failed cleanups.
Phase: network failures
The pod runs but can't reach peers / can't expose IB / can't egress.
CNI issues
$ kubectl exec <pod> -- ip a
$ kubectl exec <pod> -- ping <other-pod-ip>
$ kubectl logs -n kube-system <cni-pod-on-this-node>
If the pod has no IP, CNI failed at sandbox setup. If it has an IP but can't reach other pods, you might have:
- NetworkPolicy denies (
kubectl get networkpolicies -A). - CNI routing issue (look at CNI agent logs).
- Underlay routing issue (host can't reach other host).
IB device-plugin
For multi-node GPU jobs over RDMA, the pod needs /dev/infiniband mounted. There are two patterns:
- HostPath mount — old style; pod spec has
volumeMountsfor/dev/infiniband. rdma/<device>resource via the Network Operator IB device-plugin — modern; pod requestsrdma/hca: 1and the plugin injects the right device files.
If kubectl exec <pod> -- ls /dev/infiniband shows nothing, the device plugin isn't injecting (or hostpath isn't mounted, depending on which pattern you use).
# verify the IB device-plugin is exposing resources
$ kubectl describe node <node> | grep rdma
rdma/rdma_shared_device_a: 1
If the resource isn't advertised, the device-plugin DaemonSet isn't running on that node or isn't healthy. kubectl logs -n network-operator <ib-device-plugin-pod>.
NetworkPolicy
$ kubectl get networkpolicy -A
$ kubectl describe networkpolicy <name> -n <ns>
A default-deny policy that doesn't allow egress to other pods will block NCCL bootstrap. The fix is usually to add an egress rule allowing intra-namespace traffic, or to allow the bootstrap port range explicitly.
When kubectl exec doesn't work — debug containers
If the main container is in a state where exec is rejected (e.g., distroless image with no shell, or container doesn't expose /bin/sh), use ephemeral debug:
$ kubectl debug -it <pod> -n <ns> --image=nicolaka/netshoot --target=<container>
This attaches a netshoot container into the same PID and network namespace as the target, so you can run nsenter-style debugging without a shell in the original image:
# inside the debug container, you can see the target's processes
$ ps -ef
$ ss -tlnp # what's listening
$ ip a # network namespace
$ cat /proc/1/environ | tr '\0' '\n' # env vars of the target
For a node-level debug:
$ kubectl debug node/<node> -it --image=ubuntu
# you get a privileged pod with /host bind-mounted
$ chroot /host
$ # now you're effectively SSHed into the node
Quick playbook
| Pod state | First thing to look at |
|---|---|
Pending | kubectl describe pod → Events → FailedScheduling reason |
ContainerCreating | kubectl describe pod → Events → recent volume / network / runtime errors |
Init:CrashLoopBackOff | kubectl logs <pod> -c <init-container> |
CrashLoopBackOff | kubectl logs <pod> --previous; if GPU-related, run nvidia-smi inside |
OOMKilled | kubectl describe pod, then check resources.limits.memory and actual usage |
Evicted | kubectl describe node → conditions; clean up disk or memory pressure |
Running but not working | kubectl logs, kubectl exec ... nvidia-smi, kubectl debug |
See also
- Triage decision tree
- GPU Operator
- Network Operator
- Reservations
- ArgoCD
- MIG
- NCCL multi-node failures
- driver / firmware mismatch
External:
- K8s troubleshooting: kubernetes.io/docs/tasks/debug/
- kubectl debug docs: kubernetes.io/docs/tasks/debug/debug-application/debug-running-pod/
- CDI spec: github.com/cncf-tags/container-device-interface