Inside a training step

Connect forward, backward, synchronization, and optimizer work to infrastructure signals.

▶Practice supported commands in the command emulator — type help for the full list, or solutions for reference procedures. Simulated outputs do not validate your environment.

A slow step is a useful observation, but it is not a diagnosis. Split the step into work that runs locally and work that waits on something else. Then check which part grew relative to the same workload's baseline.

Follow the dependencies

PhaseMain workWhat to inspect
InputPrepare and transfer the next batchBatch wait, host memory, CPU, storage
ForwardCompute predictions and lossKernel timing, activations, GPU clocks
BackwardCompute gradientsKernel timing, peak memory, recomputation
SynchronizationExchange or reduce distributed stateRank arrival times, collective duration
OptimizerApply the parameter updateMemory bandwidth, state placement

This is a dependency map, not a promise of serial execution. Data preparation can overlap GPU work; gradient communication can overlap backward computation. Adding overlapping profiler durations overestimates elapsed step time.

Budget memory with assumptions

For an illustrative 3-billion-parameter model, assume BF16 weights and gradients, two FP32 Adam moments, and an FP32 master copy of weights:

StateBytes per parameterDecimal GB
Weights26
Gradients26
Adam moments824
Master weights412
Total1648

Activations, temporary buffers, communication buffers, and allocator headroom are additional. Some implementations omit master weights or store gradients differently. Sharding changes the per-rank budget; dividing 48 GB by the GPU count alone is not a capacity plan.

Investigate the phase that changed

Longer forward and backward kernels suggest a different input shape, precision path, clock state, or kernel choice. Longer exposed synchronization can mean slower transport or a peer arriving late. A timeout reported by one rank can originate in another rank's earlier failure.

Use tokens and batches to normalize the measurement and parallelism to understand which ranks exchange state.

Practice in the emulator

Open the interactive training lab. Investigate a fictional job with synthetic observations, one command at a time.