Out-of-band management evidence
Use the management path to distinguish access loss from host failure.
help for the full list, or solutions for reference procedures. Simulated outputs do not validate your environment.The out-of-band network exists so that platform evidence remains available when the host operating system or its normal network does not. Losing that access is a different observation from losing the running workload.
Check which path failed
| Observation | First comparison |
|---|---|
| BMC unreachable, job still progressing | Management routing, switch, ACL, BMC health |
| Host unreachable, BMC reachable | Power state, console, boot and hardware events |
| Several devices disappear together | Shared switch, power domain, change timeline |
| Link events coincide with host restart | Host boot identity, BMC events, scheduler history |
These comparisons narrow the scope; none proves the cause alone. A missing BMC event may reflect retention or clock problems.
Discover Redfish resources
Start at the service root and follow advertised resource links. Systems, Chassis, Managers, and LogServices are separate resources. Their member names vary across implementations; do not hardcode a vendor's system identifier into a generic collector.
Keep event timestamps and collection timestamps. If the controller clock differs from the host clock, document the offset before correlating incidents.
DMTF publishes the Redfish specifications and schemas. Use the schema and implementation documentation when selecting properties.
Capture evidence before recovery actions
Power cycling and controller resets change state and can erase useful context. First collect relevant health, power, inventory, and event information. Choose a recovery action only after identifying the failing boundary and its effect on the job.
See DPUs and BMCs and evidence collection.