Training infrastructure
Understand the training loop, identify its bottleneck, and follow the relevant infrastructure reference.
help for the full list, or solutions for reference procedures. Simulated outputs do not validate your environment.A training job connects compute, communication, input loading, and recovery into one repeating loop. Start with where that loop waits, then use the component reference to investigate the cause.
For an operational task, start with the operator workflow. The reading order below establishes the workload context before you investigate a component.
Training workflow
This section explains how the workload behaves. Hardware, network, storage, and operational procedures live in their respective sections of the reference.
| Order and guide | Question it answers | Record before moving on |
|---|---|---|
| 1. Workloads | How do training, adaptation, and serving differ? | Workload type and completion target |
| 2. The training step | Where do compute and communication occur? | Compute and communication phases |
| 3. Tokens, batches, steps | How is useful progress counted? | Batch definition and progress metric |
| 4. Parallelism and placement | How is model state and work divided across ranks? | Rank groups and placement constraints |
| 5. Data loading | What must happen before the next batch is ready? | Batch wait and input path |
| 6. Checkpointing | How does a job save state and recover? | Save and tested restore requirements |
| 7. Cost per completed workload | Does an improvement reduce total completion cost? | Completion time and resource cost |
Find the responsible component
Each topic has one reference page. Follow the symptom into that page instead of repeating its setup and commands here.
| Observation | Reference | Evidence to keep |
|---|---|---|
| One node runs slower than its peers | NVIDIA Stack: GPU data path | Clocks, temperatures, kernel durations, placement |
| Local communication is slow | Networking: NVLink / NVSwitch | Domain boundary, peer paths, link counters |
| One node is fast; several are slow | Networking: NCCL | Message sizes, rank arrival, transport |
| Performance changes with placement | Networking: topology | GPU-to-NIC locality and shared-link capacity |
| Input reads and saves compete | Storage: training workloads | Read/write demand, metadata latency, cache state |
| A link is active but the job regressed | Operations: training fabric triage | Negotiated rate, endpoint identity, counter deltas |
| Host access disappears | Operations: out-of-band management | Power state, boot history, hardware events |
| Evidence needs to be shared | Operations: diagnostic snapshot | Scope, UTC interval, raw output, missing data |
The command emulator includes fictional labs for input wait, multi-node slowdown, and a late rank. Numerical examples state their assumptions; use measurements from your own workload to establish a baseline.
System context
Conceptual illustration. Connections show system roles, not a physical cabling plan.