Operator workflow

Choose a task, collect the required context, and finish with a result another operator can verify.

▶Practice supported commands in the command emulator — type help for the full list, or solutions for reference procedures. Simulated outputs do not validate your environment.

Use this page to choose a route through the reference. Follow the stages for a new environment; for an existing incident, go directly to Diagnose and recover. The sidebar groups the detailed references by system.

Choose your task

Your taskStart hereFinish with
Understand a training workloadUnderstand the workloadWorkload requirements and a progress metric
Prepare a node or clusterPrepare the environmentRecorded versions, topology, and scheduler configuration
Accept a new or changed environmentValidate before handoffReproducible baseline and an explicit acceptance decision
Run the service day to dayOperate and monitorHealth checks, alerts, and a tested recovery path
Investigate a failure or slowdownDiagnose and recoverEvidence, a controlled comparison, and a verified outcome

Understand the workload

Before you start: identify the workload owner, model or application, GPU count, target completion time, and recovery requirements.

  1. Read training, adaptation, and serving to identify the workload type.
  2. Follow the training workflow to understand steps, progress, parallelism, input loading, and checkpoints.
  3. Record memory requirements, placement constraints, input demand, and checkpoint size. Use capacity planning when sizing an allocation.

Done when: you can state the useful progress metric, the required resources, and what interruption the workload can tolerate. GPU utilization alone is not the success criterion.

Prepare the environment

Before you start: record the hardware inventory, supported software versions, intended topology, and the change window.

LayerReferenceRecord before validation
FacilityPower and coolingPower and cooling capacity for the intended load
NodeNVIDIA drivers or AMD ROCm; kernel compatibilityGPU model, firmware, kernel, driver, runtime
FabricRDMA, then InfiniBand or RoCELink layer, expected rates, GPU-to-NIC mapping, peers
StorageTraining storageClient path, read/write demand, durability requirements
SchedulingKubernetes or Slurm / SUNKAllocation, GPU binding, access and isolation rules

Choose the references matching the deployed platform. Commands and settings depend on the recorded versions and topology; review changes against that environment before applying them.

Done when: the intended configuration is documented and a test allocation is available. Installation success is not workload acceptance.

Validate before handoff

Before you start: agree on acceptance criteria, reserve test resources, and identify a comparable healthy baseline.

  1. Check the node using the health check runbook.
  2. Validate the endpoint path with RDMA perftest.
  3. Compare local and distributed collectives with NCCL tests and multi-node validation.
  4. Run a representative workload. Record input wait, step time, useful throughput, checkpoint duration, and a restore result.
  5. Preserve commands, versions, rank placement, correctness results, and timed observations in an evidence bundle.

Done when: the representative workload meets the agreed criteria and another operator can reproduce the measurement. Record exceptions and untested paths explicitly; a passing benchmark only covers the conditions tested.

Operate and monitor

Before you start: establish the accepted baseline, service owner, escalation path, and recovery procedure.

  1. Use monitoring to connect workload progress with GPU, fabric, and host telemetry.
  2. Configure alerts with an owner and an investigation path.
  3. Follow the health check runbook and, for SUNK, the day-2 runbook.
  4. Exercise checkpoint recovery and retain the measured restart time.

Done when: another operator can recognize a regression, find the relevant evidence, and follow the recovery procedure. Keep the baseline current after validated changes.

Diagnose and recover

Before you start: record impact, affected jobs and nodes, the UTC interval, recent changes, and one healthy comparison if available.

  1. Collect a bounded diagnostic snapshot.
  2. Use triage to locate the failing layer. For a training slowdown, use the symptom map.
  3. Choose the focused procedure: training fabric, NCCL failures, GPU pod failures, or host unavailable.
  4. Follow incident response for ownership, changes, communication, and recovery.
  5. Repeat the failing comparison after the correction, then verify useful workload progress and recovery state.

Done when: the original failure no longer reproduces under the same conditions, the affected workload is progressing, and the evidence records the outcome and any remaining uncertainty.

Practice and hand over

Use the command emulator to practice reading simulated outputs. Its fictional labs do not execute commands or validate a real cluster.

For a work handoff, record: scope and owner; versions and placement; commands and timestamps; expected versus observed results; changes and rollback; remaining limitations; and the next action. The runbook template provides a reusable structure.