Tokens, batches, and steps

Make throughput comparisons use the same amount of work.

▶Practice supported commands in the command emulator — type help for the full list, or solutions for reference procedures. Simulated outputs do not validate your environment.

A log reporting milliseconds per step is incomplete without the batch definition. Bigger batches and longer sequences can increase step time while improving total throughput.

Keep the units explicit

A token is an element of a tokenized sequence. A sample can contain a fixed-length sequence, a variable-length sequence, or several packed examples. Record whether a token counter includes padding.

For a conventional data-parallel setup:

global samples per optimizer step
  = samples per microbatch per replica
  × gradient accumulation steps
  × data-parallel replicas

Tensor- and pipeline-parallel ranks cooperate on the same replica; they do not multiply the data-parallel sample count.

A worked example

Assume 4 data-parallel replicas, 2 sequences per microbatch, 8 accumulation steps, and 2,048 non-padding tokens per sequence:

sequences per optimizer step = 4 × 2 × 8 = 64
tokens per optimizer step    = 64 × 2,048 = 131,072
at 2 seconds per step       = 65,536 tokens/second

With variable sequence lengths, sum actual tokens instead of multiplying by a fixed length. If logs use “step” for a microbatch rather than an optimizer update, reconcile that before calculating throughput.

Separate startup from steady state

Compilation, cache warming, worker startup, and memory allocation can make early steps atypical. Report the warmup interval and the measurement window. Keep tail latency as well as the median: occasional long steps may dominate a short run.

A larger global batch can change optimization behavior. A throughput comparison is not automatically a comparison of time to the same model quality.

See data loading and the training step.