Congestion and shared paths

Correlate queueing with the traffic that creates it.

▶Practice supported commands in the command emulator — type help for the full list, or solutions for reference procedures. Simulated outputs do not validate your environment.

Congestion appears when traffic arrives faster than a shared resource can serve it. GPU fabrics and storage networks can experience it at different times and through different flow-control mechanisms.

Identify the traffic pattern

PatternTypical triggerEvidence
IncastMany workers write toward a shared destinationReceiver load, queueing, write latency
Sustained contentionJobs share an uplink or storage poolThroughput changes with competing load
Bursty communicationSynchronized training phasesCollective timing aligned with port metrics
Input starvationReads compete with checkpoint writesBatch wait rises during saves

InfiniBand uses link-level credit flow control. RoCE deployments may use Ethernet mechanisms such as ECN and PFC. Interpret counters and behavior for the actual fabric; do not transfer tuning values between them.

Test the correlation

Plot step time, checkpoint activity, read latency, and relevant network counters on one clock. If the same pauses recur during saves, repeat a controlled workload with a changed save schedule or isolated destination.

A correlation narrows the investigation; it does not establish which queue is responsible. Storage service time and host preprocessing can produce similar symptoms.

Change one constraint at a time

Potential remedies include staggering saves, limiting background traffic, improving placement, or adding capacity at the measured bottleneck. More buffering alone cannot solve a sustained capacity deficit.

Verify that the change improves completed work and tail latency. A lower network utilization graph can simply mean that the application is doing less work.

See storage for training, data loading, and RoCE.