Out-of-band management evidence

Use the management path to distinguish access loss from host failure.

▶Practice supported commands in the command emulator — type help for the full list, or solutions for reference procedures. Simulated outputs do not validate your environment.

The out-of-band network exists so that platform evidence remains available when the host operating system or its normal network does not. Losing that access is a different observation from losing the running workload.

Check which path failed

ObservationFirst comparison
BMC unreachable, job still progressingManagement routing, switch, ACL, BMC health
Host unreachable, BMC reachablePower state, console, boot and hardware events
Several devices disappear togetherShared switch, power domain, change timeline
Link events coincide with host restartHost boot identity, BMC events, scheduler history

These comparisons narrow the scope; none proves the cause alone. A missing BMC event may reflect retention or clock problems.

Discover Redfish resources

Start at the service root and follow advertised resource links. Systems, Chassis, Managers, and LogServices are separate resources. Their member names vary across implementations; do not hardcode a vendor's system identifier into a generic collector.

Keep event timestamps and collection timestamps. If the controller clock differs from the host clock, document the offset before correlating incidents.

DMTF publishes the Redfish specifications and schemas. Use the schema and implementation documentation when selecting properties.

Capture evidence before recovery actions

Power cycling and controller resets change state and can erase useful context. First collect relevant health, power, inventory, and event information. Choose a recovery action only after identifying the failing boundary and its effect on the job.

See DPUs and BMCs and evidence collection.