Training, adaptation, and serving
Identify the workload before choosing a performance metric.
help for the full list, or solutions for reference procedures. Simulated outputs do not validate your environment.GPU activity alone cannot tell you whether a job is healthy. A training run measures progress toward a model; a serving system measures how quickly requests finish. Even on identical hardware, those workloads put pressure on different paths.
What the workload asks of the cluster
| Workload | Useful progress metric | Common pressure |
|---|---|---|
| Pre-training | Tokens processed at a fixed training configuration | Compute, collective communication, recovery |
| Full fine-tuning | Examples or tokens per optimizer step | Model state, activations, data preparation |
| Adapter tuning | Task progress with a small trainable parameter set | Base-model memory, activations, input pipeline |
| Inference | Request throughput and latency distribution | Weight reads, KV cache, batching, interconnect |
These are starting points, not hardware prescriptions. Model-parallel inference can communicate on every layer. Fine-tuning may occupy one GPU or many nodes. Mixture-of-experts models add communication patterns that a dense-model example does not capture.
Read a run as a timeline
Allocation, initialization, first-batch preparation, steady-state steps, checkpoint saves, and recovery have different costs. Measure them separately. A short run can be dominated by startup even when its steady-state throughput is excellent.
Synchronous training couples ranks through communication. A rank failure can abort its communicator and force a restart; the exact recovery boundary depends on the framework. Elastic launchers can restart workers, but they do not make lost progress free.
Make a useful comparison
Hold the model, sequence length, precision, batch definition, and software environment constant. Record elapsed time and completed work, then compare the slow run with a known-good run at the same scale. A larger allocation that finishes only slightly sooner may be a worse use of capacity.
Continue with the training step, parallelism, and inference versus training operations.