Parallelism and placement

Map each communication group to the hardware paths it uses.

▶Practice supported commands in the command emulator — type help for the full list, or solutions for reference procedures. Simulated outputs do not validate your environment.

Adding GPUs can provide memory capacity, throughput, or both. The communication cost depends on how the model and its state are divided, not just on GPU count.

Identify the group

StrategyWhat each rank ownsCommunication to expect
Data parallelA model replica, a different batchGradient reduction between replicas
Sharded data parallelA portion of model or optimizer stateGathers and reductions according to the sharding scheme
Tensor parallelA slice of layer computationFrequent exchanges within layers
Pipeline parallelA set of layersActivations forward, gradients backward
Expert parallelA subset of expertsToken dispatch and result exchange

Real jobs combine strategies. Always record group membership; a “64-GPU run” does not say which GPUs communicate with one another.

Place the most frequent exchanges deliberately

For an illustrative 16-GPU job, TP=4 and DP=4 gives four replicas, each spread across four GPUs. Keeping each tensor-parallel group within a fast local domain can reduce exposed layer-to-layer communication. The four replicas still need their data-parallel synchronization path.

This is a placement example, not a recommendation for every model. Available memory, layer dimensions, kernel support, and network topology constrain valid choices.

Account for pipeline bubbles

In a simplified non-interleaved pipeline with p stages and m microbatches, the bubble fraction is approximately (p − 1) / (m + p − 1). With 4 stages and 12 microbatches, that is 3/15 = 20%. Real schedules, imbalanced stages, and communication change the result.

Increasing microbatches may improve pipeline occupancy but changes activation storage and scheduling overhead. Inspect the timeline before assuming another parallelism dimension will help.

See NVLink domains, collectives, and batch arithmetic.