Read the GPU data path

Locate compute, memory, and host-transfer bottlenecks before changing the fabric.

▶Practice supported commands in the command emulator — type help for the full list, or solutions for reference procedures. Simulated outputs do not validate your environment.

A GPU sits on several paths with different limits. Tensor operations use execution units; model state and activations occupy HBM; host transfers traverse the system's PCIe and CPU topology. A high utilization percentage does not identify which resource is limiting progress.

Map a node before comparing it

nvidia-smi -L
nvidia-smi topo -m
nvidia-smi -q

Keep device UUIDs and bus IDs with the output. Device indices can change across environments, so “GPU 0” is not a durable inventory identifier.

ResourceEvidenceInterpretation
ExecutionKernel durations, clocks, power limitsLonger kernels need workload and clock context
HBM capacityAllocated/reserved memory, peak useOut-of-memory can be a workload sizing issue
HBM bandwidthProfiler memory throughputCapacity and bandwidth are different constraints
PCIeNegotiated generation/width under loadCompare with the platform's supported topology
Local interconnectLink state and peer-transfer behaviorCheck the path actually used by the workload

Compare peers under the same work

If one rank runs a longer kernel, compare temperature, clocks, power policy, input shape, and process placement at the same timestamp. A lower idle link speed is not, by itself, proof of a degraded PCIe link.

Record the first relevant GPU error with its device identity and timestamp. Interpret it using NVIDIA's Xid documentation; a code alone is not a replacement decision.

For hardware-specific capacity and bandwidth, use the exact SKU's specification. See GPU generations, DCGM, and PCIe topology.