Cost per completed workload

Separate allocation cost, useful throughput, and recovery overhead.

▶Practice supported commands in the command emulator — type help for the full list, or solutions for reference procedures. Simulated outputs do not validate your environment.

A cheaper GPU-hour does not guarantee a cheaper completed run. Compare the amount of useful work finished for the total allocation time, including startup, stalled steps, checkpoint pauses, and replay.

Start with an explicit cost model

allocation cost = GPU count × elapsed hours × assumed rate
cost per million tokens = allocation cost / completed tokens × 1,000,000

For a hypothetical 32-GPU allocation at $2 per GPU-hour, 10 hours costs $640. Completing the same workload in 8 hours costs $512: a $128 reduction. These are arithmetic inputs, not provider quotes; storage, networking, reservations, and other charges are outside this example.

If throughput falls by 20%, a fixed amount of work takes 1 / 0.8 = 1.25 times as long. The runtime and allocation cost rise by 25%, assuming the rate and GPU count stay fixed.

Recovery consumes allocation time too

A failure can add detection, replacement, initialization, checkpoint loading, and replay. Count the allocation that remains reserved during each phase. Do not bill the same interval twice when diagnosis and replacement overlap.

With evenly distributed failures and a fixed checkpoint interval, average lost progress is approximately half the interval. The worst case is nearly a whole interval. Actual loss depends on which checkpoint is complete and recoverable.

Use MFU as supporting evidence

Model FLOPS utilization compares estimated useful model computation per second with the selected hardware peak. The estimate and denominator depend on model accounting, precision, GPU SKU, and whether sparsity is assumed. Compare compatible measurements; there is no single healthy percentage for every job.

A throughput improvement matters only if the training configuration and result remain comparable. Keep step time, completed tokens, failure count, and restore duration next to utilization.

See checkpointing and capacity planning.