Training fabric troubleshooting

Use controlled comparisons to locate the failing boundary.

▶Practice supported commands in the command emulator — type help for the full list, or solutions for reference procedures. Simulated outputs do not validate your environment.

Start with the first change in useful progress, not the loudest final alarm. Capture the job's placement and earliest relevant error before testing a hypothesis.

Choose the smallest useful comparison

SymptomFirst comparisonWhat it narrows
Single-node run is slowSame workload on a known-good nodeLocal compute, memory, input path
Local run is fast, distributed run is slowFixed two-node benchmarkTransport or inter-node path
Only one pair is slowSwap one endpoint at a timeEndpoint versus shared route
Periodic stallsSave timeline and input waitCheckpoint/input contention
Management disappears, training continuesSeparate data and management probesAccess path versus workload path

An isolated benchmark passing does not prove the full job is healthy. Scale, shared load, message pattern, and application behavior can expose different failures.

Check negotiated width and speed, both endpoint identities, and error-counter deltas. Compare the intended cable map with live peers. Then repeat the same measurement against a healthy path.

Do not recable based on a hostname pattern alone. Resolve the physical connectors and any breakout mapping first. See fabric inventory.

Collective timeout

Read logs from all ranks. Look for an earlier out-of-memory event, process exit, GPU error, or host interruption before assuming a fabric fault. Confirm collective order and rank membership with the application owner.

If the evidence points to transport, use the NCCL isolation sequence with fixed placement and versions.

Several endpoints change together

Group events by time and shared dependency: host, switch, power, or management path. Correlate platform uptime and scheduler events. Host-facing link transitions may follow a restart; they do not prove what caused that restart.

Keep out-of-band evidence and the workload timeline together.

Close with a repeatable result

After a corrective action, repeat the failing comparison at the same scale. Verify workload progress, error-counter behavior, and recovery state. Record any remaining limitation rather than declaring the entire fabric healthy from one successful test.

See multi-node validation and incident response.