Latency on the critical path

Distinguish small-message overhead, transfer time, and a late peer.

▶Practice supported commands in the command emulator — type help for the full list, or solutions for reference procedures. Simulated outputs do not validate your environment.

Low latency matters when the workload repeatedly waits for small exchanges. Large transfers can be dominated by bandwidth instead. A single hardware latency number cannot predict end-to-end step time.

Use a simple model

transfer time ≈ startup overhead + payload / effective bandwidth

As an illustrative calculation, 4 KB at 25 GB/s takes about 0.16 microseconds for the payload alone. A 2-microsecond startup cost dominates it. At 256 MB, the payload term is about 10.24 milliseconds, and the same startup cost is comparatively small. These inputs are hypothetical, not measured platform specifications.

A collective adds topology, synchronization, and algorithm effects. Do not apply the point-to-point estimate as an exact collective prediction.

Measure the delay that training exposes

Communication hidden behind compute may not extend the step. The relevant time is the part still on the critical path. Likewise, a rank waiting inside NCCL may have arrived early while another rank is still computing.

Keep rank arrival timestamps and kernel timing alongside network measurements. Without that context, a compute straggler can look like a slow link.

Diagnose by message range

ObservationNext comparison
Small messages regressLaunch overhead, synchronization, CPU scheduling
Large messages regressEffective bandwidth, path, contention
Only selected pairs regressLocality, endpoint state, shared links
All peers wait on one rankWork before the collective on that rank

Use collective measurements and topology to narrow the cause before changing communication settings.