Diagnostic snapshot and evidence bundle
Collect host, GPU, fabric, and platform evidence with reproducible context.
help for the full list, or solutions for reference procedures. Simulated outputs do not validate your environment.A useful evidence bundle lets another person reproduce your conclusion without querying the live system again. Collect enough context to identify the affected device and time window, while keeping the collection bounded.
Define the question before the query
Write down the symptom, affected allocation, UTC interval, and one healthy comparison target. Decide whether you need host logs, GPU state, fabric events, storage metrics, or platform evidence.
Prefer a role that cannot modify the service. Verify each endpoint's documented behavior; an HTTP method alone is not an access-control policy.
Keep four parts
| Part | Contents |
|---|---|
| Context | Job ID, host/GPU/NIC identity, software versions |
| Raw evidence | Original responses and command output |
| Collection record | Start/end time, command or endpoint, exit/status code |
| Interpretation | Observations, hypotheses, and missing evidence |
Validate response status and structure. Handle pagination explicitly. Mark cached or incomplete results so they cannot be mistaken for a current complete snapshot.
Bound the collection
Use timeouts, limited concurrency, and a small retry budget. Stop repeated authentication failures. Broad discovery or frequent polling can burden a management service even when requests do not change configuration.
Keep credentials out of command lines and captured output. Use approved credential handling, verify TLS, and restrict access to bundles that contain internal inventory or logs. Redact before sharing outside the intended audience.
Finish with a concise observation: what changed, when, where, and how it differs from the healthy comparison. Keep the proposed cause separate until the evidence supports it.
Initial diagnostic snapshot
Use these commands to collect an initial snapshot from an authorized host. Availability and permissions depend on the installed driver and diagnostic packages. Record UTC time, host identity, and the affected job alongside the output.
Host and GPU
date -u
hostname
nvidia-smi -L
nvidia-smi topo -m
nvidia-smi -q
journalctl -k --since "15 minutes ago" --utc --no-pager
Keep UUIDs and bus IDs. Match errors to the correct GPU and time window. A current snapshot can miss a transient that occurred before collection.
Network and host pressure
ibstat
ibv_devinfo
ip -s link
vmstat 1 5
| Output | Look for | Avoid concluding |
|---|---|---|
| GPU query | Memory, clocks, power, thermal context | High utilization means useful progress |
| Topology matrix | GPU/NIC locality and peer paths | A topology label is a bandwidth result |
| HCA state | Link layer, port state, negotiated rate | Active means full performance |
| Interface counters | Timed deltas and matching interface | Historical errors caused this incident |
| Host pressure | CPU, runnable tasks, memory and I/O wait | I/O wait uniquely identifies storage |
Collectives and platform evidence
Capture NCCL INFO logs for a bounded diagnostic run. Keep rank mapping and the first errors from all ranks. For active benchmarks and interpretation, use nccl-tests.
For platform events, use the controller's documented read endpoints and discover Redfish resource links. Keep resetting counters, updating firmware, and power actions out of the initial snapshot: each changes the evidence being collected.
For continuous telemetry use DCGM. For fabric and platform context, see UFM, InfiniBand, and out-of-band management.