Training fabric troubleshooting
Use controlled comparisons to locate the failing boundary.
help for the full list, or solutions for reference procedures. Simulated outputs do not validate your environment.Start with the first change in useful progress, not the loudest final alarm. Capture the job's placement and earliest relevant error before testing a hypothesis.
Choose the smallest useful comparison
| Symptom | First comparison | What it narrows |
|---|---|---|
| Single-node run is slow | Same workload on a known-good node | Local compute, memory, input path |
| Local run is fast, distributed run is slow | Fixed two-node benchmark | Transport or inter-node path |
| Only one pair is slow | Swap one endpoint at a time | Endpoint versus shared route |
| Periodic stalls | Save timeline and input wait | Checkpoint/input contention |
| Management disappears, training continues | Separate data and management probes | Access path versus workload path |
An isolated benchmark passing does not prove the full job is healthy. Scale, shared load, message pattern, and application behavior can expose different failures.
Active link, unexpected performance
Check negotiated width and speed, both endpoint identities, and error-counter deltas. Compare the intended cable map with live peers. Then repeat the same measurement against a healthy path.
Do not recable based on a hostname pattern alone. Resolve the physical connectors and any breakout mapping first. See fabric inventory.
Collective timeout
Read logs from all ranks. Look for an earlier out-of-memory event, process exit, GPU error, or host interruption before assuming a fabric fault. Confirm collective order and rank membership with the application owner.
If the evidence points to transport, use the NCCL isolation sequence with fixed placement and versions.
Several endpoints change together
Group events by time and shared dependency: host, switch, power, or management path. Correlate platform uptime and scheduler events. Host-facing link transitions may follow a restart; they do not prove what caused that restart.
Keep out-of-band evidence and the workload timeline together.
Close with a repeatable result
After a corrective action, repeat the failing comparison at the same scale. Verify workload progress, error-counter behavior, and recovery state. Record any remaining limitation rather than declaring the entire fabric healthy from one successful test.
See multi-node validation and incident response.