Training, adaptation, and serving

Identify the workload before choosing a performance metric.

▶Practice supported commands in the command emulator — type help for the full list, or solutions for reference procedures. Simulated outputs do not validate your environment.

GPU activity alone cannot tell you whether a job is healthy. A training run measures progress toward a model; a serving system measures how quickly requests finish. Even on identical hardware, those workloads put pressure on different paths.

What the workload asks of the cluster

WorkloadUseful progress metricCommon pressure
Pre-trainingTokens processed at a fixed training configurationCompute, collective communication, recovery
Full fine-tuningExamples or tokens per optimizer stepModel state, activations, data preparation
Adapter tuningTask progress with a small trainable parameter setBase-model memory, activations, input pipeline
InferenceRequest throughput and latency distributionWeight reads, KV cache, batching, interconnect

These are starting points, not hardware prescriptions. Model-parallel inference can communicate on every layer. Fine-tuning may occupy one GPU or many nodes. Mixture-of-experts models add communication patterns that a dense-model example does not capture.

Read a run as a timeline

Allocation, initialization, first-batch preparation, steady-state steps, checkpoint saves, and recovery have different costs. Measure them separately. A short run can be dominated by startup even when its steady-state throughput is excellent.

Synchronous training couples ranks through communication. A rank failure can abort its communicator and force a restart; the exact recovery boundary depends on the framework. Elastic launchers can restart workers, but they do not make lost progress free.

Make a useful comparison

Hold the model, sequence length, precision, batch definition, and software environment constant. Record elapsed time and completed work, then compare the slow run with a known-good run at the same scale. A larger allocation that finishes only slightly sooner may be a worse use of capacity.

Continue with the training step, parallelism, and inference versus training operations.