Parallelism and placement
Map each communication group to the hardware paths it uses.
help for the full list, or solutions for reference procedures. Simulated outputs do not validate your environment.Adding GPUs can provide memory capacity, throughput, or both. The communication cost depends on how the model and its state are divided, not just on GPU count.
Identify the group
| Strategy | What each rank owns | Communication to expect |
|---|---|---|
| Data parallel | A model replica, a different batch | Gradient reduction between replicas |
| Sharded data parallel | A portion of model or optimizer state | Gathers and reductions according to the sharding scheme |
| Tensor parallel | A slice of layer computation | Frequent exchanges within layers |
| Pipeline parallel | A set of layers | Activations forward, gradients backward |
| Expert parallel | A subset of experts | Token dispatch and result exchange |
Real jobs combine strategies. Always record group membership; a “64-GPU run” does not say which GPUs communicate with one another.
Place the most frequent exchanges deliberately
For an illustrative 16-GPU job, TP=4 and DP=4 gives four replicas, each spread across four GPUs. Keeping each tensor-parallel group within a fast local domain can reduce exposed layer-to-layer communication. The four replicas still need their data-parallel synchronization path.
This is a placement example, not a recommendation for every model. Available memory, layer dimensions, kernel support, and network topology constrain valid choices.
Account for pipeline bubbles
In a simplified non-interleaved pipeline with p stages and m microbatches, the bubble fraction is approximately (p − 1) / (m + p − 1). With 4 stages and 12 microbatches, that is 3/15 = 20%. Real schedules, imbalanced stages, and communication change the result.
Increasing microbatches may improve pipeline occupancy but changes activation storage and scheduling overhead. Inspect the timeline before assuming another parallelism dimension will help.
See NVLink domains, collectives, and batch arithmetic.