Read the GPU data path
Locate compute, memory, and host-transfer bottlenecks before changing the fabric.
help for the full list, or solutions for reference procedures. Simulated outputs do not validate your environment.A GPU sits on several paths with different limits. Tensor operations use execution units; model state and activations occupy HBM; host transfers traverse the system's PCIe and CPU topology. A high utilization percentage does not identify which resource is limiting progress.
Map a node before comparing it
nvidia-smi -L
nvidia-smi topo -m
nvidia-smi -q
Keep device UUIDs and bus IDs with the output. Device indices can change across environments, so “GPU 0” is not a durable inventory identifier.
| Resource | Evidence | Interpretation |
|---|---|---|
| Execution | Kernel durations, clocks, power limits | Longer kernels need workload and clock context |
| HBM capacity | Allocated/reserved memory, peak use | Out-of-memory can be a workload sizing issue |
| HBM bandwidth | Profiler memory throughput | Capacity and bandwidth are different constraints |
| PCIe | Negotiated generation/width under load | Compare with the platform's supported topology |
| Local interconnect | Link state and peer-transfer behavior | Check the path actually used by the workload |
Compare peers under the same work
If one rank runs a longer kernel, compare temperature, clocks, power policy, input shape, and process placement at the same timestamp. A lower idle link speed is not, by itself, proof of a degraded PCIe link.
Record the first relevant GPU error with its device identity and timestamp. Interpret it using NVIDIA's Xid documentation; a code alone is not a replacement decision.
For hardware-specific capacity and bandwidth, use the exact SKU's specification. See GPU generations, DCGM, and PCIe topology.