Latency on the critical path
Distinguish small-message overhead, transfer time, and a late peer.
help for the full list, or solutions for reference procedures. Simulated outputs do not validate your environment.Low latency matters when the workload repeatedly waits for small exchanges. Large transfers can be dominated by bandwidth instead. A single hardware latency number cannot predict end-to-end step time.
Use a simple model
transfer time ≈ startup overhead + payload / effective bandwidth
As an illustrative calculation, 4 KB at 25 GB/s takes about 0.16 microseconds for the payload alone. A 2-microsecond startup cost dominates it. At 256 MB, the payload term is about 10.24 milliseconds, and the same startup cost is comparatively small. These inputs are hypothetical, not measured platform specifications.
A collective adds topology, synchronization, and algorithm effects. Do not apply the point-to-point estimate as an exact collective prediction.
Measure the delay that training exposes
Communication hidden behind compute may not extend the step. The relevant time is the part still on the critical path. Likewise, a rank waiting inside NCCL may have arrived early while another rank is still computing.
Keep rank arrival timestamps and kernel timing alongside network measurements. Without that context, a compute straggler can look like a slow link.
Diagnose by message range
| Observation | Next comparison |
|---|---|
| Small messages regress | Launch overhead, synchronization, CPU scheduling |
| Large messages regress | Effective bandwidth, path, contention |
| Only selected pairs regress | Locality, endpoint state, shared links |
| All peers wait on one rank | Work before the collective on that rank |
Use collective measurements and topology to narrow the cause before changing communication settings.