Inside a training step
Connect forward, backward, synchronization, and optimizer work to infrastructure signals.
help for the full list, or solutions for reference procedures. Simulated outputs do not validate your environment.A slow step is a useful observation, but it is not a diagnosis. Split the step into work that runs locally and work that waits on something else. Then check which part grew relative to the same workload's baseline.
Follow the dependencies
| Phase | Main work | What to inspect |
|---|---|---|
| Input | Prepare and transfer the next batch | Batch wait, host memory, CPU, storage |
| Forward | Compute predictions and loss | Kernel timing, activations, GPU clocks |
| Backward | Compute gradients | Kernel timing, peak memory, recomputation |
| Synchronization | Exchange or reduce distributed state | Rank arrival times, collective duration |
| Optimizer | Apply the parameter update | Memory bandwidth, state placement |
This is a dependency map, not a promise of serial execution. Data preparation can overlap GPU work; gradient communication can overlap backward computation. Adding overlapping profiler durations overestimates elapsed step time.
Budget memory with assumptions
For an illustrative 3-billion-parameter model, assume BF16 weights and gradients, two FP32 Adam moments, and an FP32 master copy of weights:
| State | Bytes per parameter | Decimal GB |
|---|---|---|
| Weights | 2 | 6 |
| Gradients | 2 | 6 |
| Adam moments | 8 | 24 |
| Master weights | 4 | 12 |
| Total | 16 | 48 |
Activations, temporary buffers, communication buffers, and allocator headroom are additional. Some implementations omit master weights or store gradients differently. Sharding changes the per-rank budget; dividing 48 GB by the GPU count alone is not a capacity plan.
Investigate the phase that changed
Longer forward and backward kernels suggest a different input shape, precision path, clock state, or kernel choice. Longer exposed synchronization can mean slower transport or a peer arriving late. A timeout reported by one rank can originate in another rank's earlier failure.
Use tokens and batches to normalize the measurement and parallelism to understand which ranks exchange state.
Practice in the emulator
Open the interactive training lab. Investigate a fictional job with synthetic observations, one command at a time.