From storage to a ready batch
Find the stage that leaves the GPU waiting for input.
help for the full list, or solutions for reference procedures. Simulated outputs do not validate your environment.The next batch travels through storage access, decoding, transforms, collation, host memory, and device transfer. A GPU idle gap can originate anywhere along that path.
The illustration follows the input path. Prefetching can overlap its stages with computation.
Measure the wait before increasing workers
| Observation | Check next |
|---|---|
| Slow first epoch, faster later epochs | Cache warming, worker startup, compilation |
| CPU busy while GPUs wait | Decode, augmentation, collation, CPU quotas |
| High read latency with little CPU work | Storage access pattern and contention |
| Host memory grows with worker count | Prefetch depth and per-worker state |
| Transfer gaps with batches already ready | Pinned memory, transfer scheduling, locality |
Increasing worker count can help CPU-bound preparation, but it can also increase memory use and storage pressure. Change worker count and prefetch depth independently.
Separate the stages experimentally
Compare representative real input with preloaded or synthetic input while keeping model settings fixed. If the gap disappears, inspect the input path; the experiment does not yet distinguish storage from preprocessing.
Measure time waiting for a batch separately from GPU kernel duration. For asynchronous device transfers, use profiler events or synchronization-aware timing; host call duration is not transfer completion time.
PyTorch's DataLoader documentation explains multiprocessing, pinning, and prefetch controls. Validate behavior for the installed version.
See storage, NUMA placement, and batch arithmetic.
Practice in the emulator
Open the interactive training lab. Investigate a fictional job with synthetic observations, one command at a time.