Training infrastructure

Understand the training loop, identify its bottleneck, and follow the relevant infrastructure reference.

▶Practice supported commands in the command emulator — type help for the full list, or solutions for reference procedures. Simulated outputs do not validate your environment.

A training job connects compute, communication, input loading, and recovery into one repeating loop. Start with where that loop waits, then use the component reference to investigate the cause.

For an operational task, start with the operator workflow. The reading order below establishes the workload context before you investigate a component.

Training workflow

This section explains how the workload behaves. Hardware, network, storage, and operational procedures live in their respective sections of the reference.

Order and guideQuestion it answersRecord before moving on
1. WorkloadsHow do training, adaptation, and serving differ?Workload type and completion target
2. The training stepWhere do compute and communication occur?Compute and communication phases
3. Tokens, batches, stepsHow is useful progress counted?Batch definition and progress metric
4. Parallelism and placementHow is model state and work divided across ranks?Rank groups and placement constraints
5. Data loadingWhat must happen before the next batch is ready?Batch wait and input path
6. CheckpointingHow does a job save state and recover?Save and tested restore requirements
7. Cost per completed workloadDoes an improvement reduce total completion cost?Completion time and resource cost

Find the responsible component

Each topic has one reference page. Follow the symptom into that page instead of repeating its setup and commands here.

ObservationReferenceEvidence to keep
One node runs slower than its peersNVIDIA Stack: GPU data pathClocks, temperatures, kernel durations, placement
Local communication is slowNetworking: NVLink / NVSwitchDomain boundary, peer paths, link counters
One node is fast; several are slowNetworking: NCCLMessage sizes, rank arrival, transport
Performance changes with placementNetworking: topologyGPU-to-NIC locality and shared-link capacity
Input reads and saves competeStorage: training workloadsRead/write demand, metadata latency, cache state
A link is active but the job regressedOperations: training fabric triageNegotiated rate, endpoint identity, counter deltas
Host access disappearsOperations: out-of-band managementPower state, boot history, hardware events
Evidence needs to be sharedOperations: diagnostic snapshotScope, UTC interval, raw output, missing data

The command emulator includes fictional labs for input wait, multi-node slowdown, and a late rank. Numerical examples state their assumptions; use measurements from your own workload to establish a baseline.

System context

Conceptual illustration. Connections show system roles, not a physical cabling plan.