Checkpointing and recovery
Size the save path and verify that a completed checkpoint can restore progress.
help for the full list, or solutions for reference procedures. Simulated outputs do not validate your environment.A checkpoint is useful only when the job can restore it. Treat save completion, durable storage, and a successful restart as separate milestones.
Decide what must survive
Resumable training commonly needs model parameters, optimizer state, scheduler state, progress counters, and random-number state. Data-loader or sampler state may also be required to reproduce data order. A weights-only export can serve inference without supporting an equivalent training restart.
Sharded checkpoints distribute state across files or objects. The metadata and completion protocol must let a reader distinguish a finished save from a partial one.
Estimate the I/O window
Suppose a job writes 240 GB and sustains 12 GB/s to its checkpoint destination. The transfer-only lower bound is 20 seconds. Serialization, synchronization, metadata operations, contention, and persistence can add time.
Measure three durations: how long training is blocked, how long the save continues in the background, and how long restoration takes. Asynchronous saving reduces exposed pause only when staging memory and background bandwidth are available.
Choose an interval with measured inputs
Short intervals reduce replay after a failure but spend more time saving. Long intervals do the reverse. Use observed interruption frequency and measured save cost; do not inherit an interval merely because it worked on another cluster.
Run a restore test into a separate allocation. Confirm the expected step, optimizer state, and continued progress. Record which checkpoint was selected and how much work was replayed.
PyTorch's distributed checkpoint documentation describes its save/load APIs and compatibility constraints. See storage sizing for the shared I/O budget.