Physical network: cables, transceivers, fiber types, and the BOM math you do before you order
DAC vs AOC vs transceiver+fiber, QSFP56 vs OSFP vs QSFP-DD, OM4 vs OS2, distance limits at 100/400/800G, BOM math for a rail-aligned fabric, cable labor estimates, and the cabling discipline that makes a 1000-cable build commission cleanly.
help for the full list, or solutions for copy-paste fix recipes.A clean InfiniBand fabric on paper is one Excel sheet. The same fabric on the floor is a thousand cables, each with a label, a color, a bend radius, a transceiver pair, and a destination port. Get the architecture right (see IB architecture) and you've solved 30% of the problem. The other 70% is the cable plant — what to order, how long, how to label, how to install without losing weeks to mis-routes. This page is that 70%.
For the upstream layout choices see rack design and cooling; for power side see power.
The three transport options at a glance
Every server-to-switch and switch-to-switch link is one of three things. The choice depends on distance and budget.
| Option | Reach (NDR/400G) | Per-link cost (typical) | Latency | Failure mode |
|---|---|---|---|---|
| Passive DAC (Direct Attach Copper) | up to ~3 m | $ | lowest (no optics) | rare; usually mechanical |
| Active DAC (with conditioning) | up to ~5-7 m | $$ | low | active components can fail |
| AOC (Active Optical Cable) | up to ~30-100 m | $$ | low | bend / kink at fiber ends |
| Transceivers + fiber patch | 100 m to 2 km+ | $$$ | low | optic itself, dirty connector |
Some operator-level intuition:
- Passive DAC up to 3 m covers in-rack and immediate-neighbor-rack links. They're cheap, low-latency, and have no powered components to fail. If you can get away with DAC, get away with DAC.
- Active DAC extends to 5-7 m at 400G but the price approaches AOC, so usage is niche.
- AOC up to 100 m covers most leaf-to-spine and end-of-row links. The cable is sold as a pre-terminated unit with optics permanently attached at both ends — you don't clean connectors, you don't swap optics, you just plug it in.
- Transceivers + fiber patch are how you build long links and how you get flexibility — the transceivers and the fiber are independent, so you can swap one without the other. More moving parts means more failure modes.
For 800G NDR-class IB / 400G Ethernet at AI-cluster scale, the dominant pattern is DAC for in-rack, AOC for in-row, transceiver+fiber for cross-row and farther. The exact 3 m / 30 m / 100 m boundaries are spec ceilings; actual deployment usually leaves margin (DAC <2.5 m, AOC <80 m).
Speeds, connector forms, and what they look like
The connector form factor evolves with speed. They all look superficially similar but they aren't interchangeable.
| Form | Lanes × rate | Aggregate | Common use |
|---|---|---|---|
| QSFP+ | 4 × 10G NRZ | 40G | legacy 40G Ethernet |
| QSFP28 | 4 × 25G NRZ | 100G | 100G EDR IB / 100G Ethernet |
| QSFP56 | 4 × 50G PAM4 | 200G | 200G HDR IB / 200G Ethernet |
| QSFP112 | 4 × 100G PAM4 | 400G | 400G NDR HCA-side / 400G Ethernet |
| OSFP | 8 × 50G PAM4 (HDR-era) | 400G | early 400G |
| OSFP | 8 × 100G PAM4 (NDR) | 800G | NDR IB switches, GB200 NICs |
| QSFP-DD | 8 × 50G or 8 × 100G PAM4 | 400G or 800G | 400/800G Ethernet, some IB |
| OSFP-XD | 16 × 100G PAM4 | 1.6T | XDR IB / 1.6T Ethernet |
Key facts that bite operators if they assume all "QSFP" plugs are the same:
- QSFP-DD is backward-compatible with QSFP/QSFP28/QSFP56 — a QSFP28 transceiver fits a QSFP-DD cage. OSFP is not backward-compatible. An OSFP optic does not fit a QSFP-DD cage.
- NVIDIA NDR switches use OSFP cages. NIC side (ConnectX-7 NDR) is QSFP112. So NDR cables are typically OSFP <-> 2× QSFP112 (one NDR switch port "splits" into two HCA links of 400G each).
- Power per optic matters at scale. An 800G OSFP can pull 13-15 W; a 400G QSFP-DD around 8-10 W. A leaf with 64 ports of 800G OSFP burns ~1 kW just in optics.
- Heat sinks: OSFP has built-in heat fins; QSFP-DD does not (the cage in the switch is the heatsink). This affects switch port power budget but is usually transparent to the operator.
Fiber types: when MMF and when SMF
Two fiber types matter for AI clusters.
Multimode (OM3 / OM4 / OM5)
- 50 μm core, large enough that multiple light modes propagate. Cheaper transceivers (VCSELs) but distance-limited.
- OM3: rated 100 m at 100GBASE-SR4. Mostly legacy now.
- OM4: rated 100 m at 400GBASE-SR4 / SR8 (PAM4 era), often deployed at shorter actual distances.
- OM5: wide-band MMF, supports more wavelengths for SWDM. Niche but used in some 400G/800G rollouts.
- Connector: typically MPO/MTP-12 or MTP-16 ferrule for parallel fiber, or LC-LC for serial.
For AI fabrics, MMF + parallel SR optics over MPO is the budget option for 30-100 m links. In practice, AOC cables replace this for distances up to 100 m because the cable + optics together is cheaper than separate optics + fiber + connector cleaning labor.
Single-mode (OS2)
- 9 μm core, single light mode. More expensive transceivers (DFB or EML lasers) but km-scale reach.
- OS2: standard SMF, used for everything from 100 m to 80 km depending on optic.
- Reach depends on the optic, not the fiber. Same OS2 fiber carries 500 m DR4, 2 km FR4, or 10 km LR4.
- Connector: LC-LC for duplex serial, or MPO-12 for parallel.
For AI build-outs, SMF is what you pull through the building's structured cabling. It future-proofs the cable plant — when you upgrade from 400G to 800G to 1.6T over five years, the fiber stays the same and you swap optics. MMF would force a fiber pull at every speed step.
The standard AI-hall pattern: pre-terminate SMF trunks (12, 24, or 48 fiber MPO bundles) between rows / aisles, land them on LC patch panels or MPO cassettes at row ends, run AOCs or short DAC/AOC inside the row.
Distance reach by transceiver type at 400G/800G
NDR-class IB and 400/800G Ethernet both use PAM4 signaling at the same line rate. Reach numbers:
| Transceiver | Fiber | Reach | Notes |
|---|---|---|---|
| 400G-SR4 | OM4 MMF (8 fiber) | 100 m | parallel, MPO-12 |
| 400G-SR8 | OM4 MMF (16 fiber) | 100 m | parallel, MPO-16 |
| 400G-DR4 | OS2 SMF | 500 m | parallel, MPO-12 |
| 400G-FR4 | OS2 SMF (duplex) | 2 km | LC-LC |
| 400G-LR4 | OS2 SMF | 10 km | LC-LC |
| 800G-SR8 / 2×SR4 | OM4 MMF | ~50 m | parallel, MPO-16 |
| 800G-DR8 / 2×DR4 | OS2 SMF | 500 m | parallel, MPO-16 |
| 800G-FR4 / 2×FR4 | OS2 SMF | 2 km | LC-LC, two breakouts |
| 800G-XDR4 | OS2 SMF | 2 km | extended DR variant |
DAC and AOC reach at the same speeds:
| Speed | Passive DAC | Active DAC | AOC |
|---|---|---|---|
| 100G | 5 m | 7 m | 100 m |
| 200G | 3 m | 5 m | 100 m |
| 400G | 2-3 m | 3-5 m | 30-50 m typical, up to 100 m |
| 800G | 1-2 m (limit) | 3 m | 30 m typical, up to 100 m for some |
The pattern: as line rate climbs, copper distance shrinks. At 100G a 5 m DAC was fine; at 800G you're at 1-2 m. This is why GB200-class deployments push everything outside an immediate cluster onto fiber.
Cabling labor: how long does this actually take
A trained two-person installer team can land roughly 40-80 cables per day with full labeling, photography, and continuity testing. That's the steady-state number for a build-out that's been planned (cable lengths pre-cut, labels pre-printed, BOM organized in delivery boxes by destination).
What that means for project schedules:
1 SU of DGX H100 SuperPOD = 32 nodes * 8 IB rails + storage + mgmt
~= 256 IB cables + 32 storage + 32 mgmt + 32 OOB + leaf-to-spine
~= 350-400 cables per SU
At 60 cables/day -> ~6-7 working days per SU for IB + storage,
plus another 2-3 days for leaf-to-spine and acceptance tests.
For 4 SU (128 nodes, ~1500 cables):
Cabling labor: ~3-4 weeks of trained installer time
Total project: 6-8 weeks including rack landing, power validation, link bring-up
Where projects fall behind:
- Cables not pre-cut to length. You're now making cables on site, slowly. Plan: order to length from a vendor with a pre-config service, label at the factory.
- Labels printed on site. A thermal printer next to the rack costs hours per label. Plan: pre-print all labels in a controlled environment, deliver in labeled bags by destination port.
- Cable trays not pre-installed. The electrician didn't finish before the network team showed up. Plan: tray install is GC scope, finished and signed off before day-1 of cabling.
- Cable count under-estimated. You ordered 1500 cables but you actually need 1650 because the BOM forgot the storage rails. Plan: order +5-10% spare on every length/type.
BOM math: from architecture to a parts list
The cable BOM falls out of the topology + rack layout. The math is straightforward but easy to get wrong by 10-20% if you skip a layer.
IB fabric cables, rail-aligned
For a rail-aligned IB fabric with N GPU nodes, R rails per node, and a leaf radix K:
Server-to-leaf cables:
count = N * R (one cable per HCA, per rail)
length: depends on whether leaves are ToR (1-3 m) or EoR (5-30 m)
Leaf-to-spine cables (full bisection):
leaf_count = ceil(N * R / (K/2)) (each leaf uses K/2 down-ports per rail)
spine_count = ceil(leaf_count * (K/2) / K)
count = leaf_count * (K/2) (each leaf uses K/2 up-ports)
length: row-to-row, typically 10-50 m
Total IB cables = N*R + leaf_count*(K/2)
Worked example: 32-node SuperPOD H100 SU, 8 rails/node, 64-port leaves:
Server-to-leaf: 32 * 8 = 256 cables
Leaf count: 256 / 32 = 8 leaves (one per rail, common in rail-aligned)
Leaf-to-spine: 8 * 32 = 256 cables
Total IB: 512 cables for 32 nodes
Per-node ratio: 512 / 32 = 16 IB cables per node when you include both halves of the fabric. This is the number of fiber endpoints touching every compute rack on the IB side alone.
Storage Ethernet, OOB, mgmt
For each compute node you also need:
- 1-2 storage links (200/400G Ethernet, AOC or fiber).
- 1 management link (1G or 10G Ethernet, copper or AOC).
- 1 OOB / BMC link (1G Ethernet, copper).
So compute side adds ~4 cables per node outside IB. For 32 nodes that's 128 more cables. Total fabric BOM for one SU including non-IB: roughly 640 cables + leaf-to-spine.
Fiber patch and transceiver count
If half your IB cables are fiber + transceivers (because they're long), each one consumes:
- 1 transceiver at each end = 2 transceivers per link.
- 1 fiber patch / trunk run, with appropriate connectors.
For 256 long IB links:
- 512 transceivers (cost dominates the BOM at scale).
- 256 fiber patches.
Transceivers are typically 60-80% of the fabric cost on AI builds. Failure rates are <1% per year per port for quality optics — but at 512 optics that's ~5 failures per year, so spares plus a return-merchandise process matter.
The cabling discipline checklist
What separates a build that commissions in two weeks from one that limps for two months.
Pre-install
- Architecture frozen (rail count, leaf radix, spine count).
- Rack elevations drawn for every rack — every U accounted for.
- Cable BOM with length + count + connector + side (A/B), 5-10% spare.
- Color-code policy for rails, A/B power, storage, mgmt — written down, posted on the wall.
- Labels printed for every cable end, organized in destination-grouped bags.
- Cable trays installed and signed off by GC.
- PDUs, busbar, racks delivered and physically positioned.
- Test gear ready: OTDR for fiber, port-by-port continuity, optical power meter.
Install
- Land power first; verify A/B feeds independently before any node enters the rack.
- Land switches second; verify management network reachability.
- Land nodes third; rack and stack with at least one operator-pair per rack.
- Cable: power first (A and B verified), then OOB, then management Ethernet, then storage, then IB.
- Per cable: label at each end as you go (don't batch — labels get mixed up).
- Per cable: bend radius respected, service loop at node end, strain relief verified.
- Photograph every rack front and rear at three checkpoints: bare rack, after power+OOB, after full install.
Acceptance
- Walk every cable visually: color matches expected rail, label at both ends, no kinks or bends below radius.
- Continuity test every link: switch port shows up, link rate is at expected speed (no degradation), error counters at zero.
- Optical power readings for every fiber link, in the expected window (typically -5 to -2 dBm for short reach, vendor-specific).
- DCGM / nvidia-smi visible on every node.
-
ibstatusshows all HCAs Active at the expected rate. - NCCL ring test on a single node (intra-node) — should pass.
- NCCL allreduce on 2 nodes - should hit ~80%+ of theoretical bandwidth.
- NCCL allreduce on full cluster — should hit the published reference number for the topology.
If anything fails, fix it before you sign off. Re-cabling after the customer has run jobs is much more expensive than re-cabling at commissioning.
Tracing tools and trace discipline
When something does fail in year two and you have to find one bad cable in a thousand:
Visual tracers
A fiber trace pen (visual fault locator, VFL) shoots a red laser down the fiber. You can see the spot at the far end. Useful for confirming "cable A really does go to port B" when you already have line of sight.
Toner / continuity testers
Plug a tester at one end, walk the suspect path with a probe, listen for the tone. Works fine for copper (DAC, Ethernet); for fiber, use the VFL or an optical loopback test.
OTDR for fiber spans
An optical time-domain reflectometer measures loss along the fiber and pinpoints where a break is, in meters. Essential for long runs and post-incident forensics.
Switch-side telemetry
ethtool / ibstat / vendor MIBs report:
- Link rate and width.
- BER counters and FEC corrections.
- Optical power readings (if the optic supports DDM/DOM).
A cable that's slowly going bad shows up as rising FEC corrections weeks before it dies. Plot these per port, alert on rate-of-change.
Connector cleanliness: the failure mode you can't see
Single-mode fiber connectors mate against a polished ferrule with a 9 μm core. A dust speck of 5 μm is enough to scatter light, raise insertion loss by 1-3 dB, and turn a working link into a flapping link. Connector cleanliness is the most common cause of "the link won't come up but everything checks out" issues.
What dirt looks like at a connector
Inspection scopes (200x video probe) reveal:
- Pristine: clean glass, no marks, no debris.
- Surface contamination: dust particles, finger oil — easily cleaned.
- Scratches: physical damage to the polish — non-recoverable, replace the connector.
- Pits: dimples in the glass surface — non-recoverable.
You inspect every connector before you mate it. Every one. Pristine install crews build this into muscle memory.
Cleaning kit
- Click-style cleaner (one-shot reel cleaner): designed for LC, MTP, etc. Two clicks per surface, dispose of the used reel section.
- Lint-free wipes + isopropyl alcohol (99%+ purity): backup for stubborn contamination.
- Inspection scope: pre- and post-clean confirmation.
When to clean
- At install: every connector before mating, no exceptions.
- At service: every connector when unplugging, before remating.
- At suspected fault: any link with elevated FEC corrections, BER, or insertion loss.
A site where nobody owns connector cleaning is a site where weird flapping links accumulate over months. Make it part of the install SOP and the maintenance SOP.
Reading a power meter
For each fiber link you should know:
- Tx power output (vendor spec, typically -2 to +2 dBm for short-reach 400G).
- Rx power expected (Tx - link loss, where link loss is fiber attenuation + connector loss).
- Acceptable Rx range (vendor spec, typically -8 to -2 dBm for short-reach).
A reading below the lower bound = link will work but with marginal SNR; FEC will work harder, errors will accumulate, performance will degrade. Above the upper bound = saturation, distortion, also bad. Most modern optics report Tx and Rx in their DDM/DOM telemetry — read it, log it, alert on drift.
Optical budget: how much loss can you afford
For short-reach 400G/800G optics, the typical link loss budget is 3-5 dB. Where it goes:
Total budget: ~4 dB typical
- Fiber attenuation: 0.4 dB/km for SMF * link_km
- Connector loss: 0.3-0.5 dB per connector pair (LC), 0.3-0.7 dB (MPO)
- Patch panel loss: 0.3-0.5 dB per pass-through
- Splice loss: 0.05-0.1 dB per fusion splice (rare in DC)
- Margin: ~1 dB
For a typical 50m intra-row link with two patch panels:
Fiber: 0.4 * 0.05 = 0.02 dB
Connectors: 3 * 0.4 = 1.2 dB (one at each end + 1 mid)
Patch panel: 2 * 0.4 = 0.8 dB
Total loss: 2.0 dB (well within 4 dB budget)
For longer runs (cross-building, campus links to 2 km), you start eating into margin:
2 km link, OS2 SMF, 2 patch panels:
Fiber: 0.4 * 2 = 0.8 dB
Connectors: 3 * 0.4 = 1.2 dB
Patch panel: 2 * 0.4 = 0.8 dB
Total: 2.8 dB
Below 4 dB FR4 budget: OK with ~1.2 dB margin.
If you have less than 0.5 dB margin, expect intermittent link issues over years.
OSFP vs QSFP-DD: which to standardize on
For a new build today, the choice between OSFP-only and QSFP-DD-only matters for inventory, spares, and operator training.
OSFP arguments
- Built-in heat sink fins = better thermal margin in dense switches.
- Dominant for 800G IB (NVIDIA Quantum-2, Quantum-X800) — if you're building IB, you're using OSFP at the switch.
- Lower per-port heat in densest switches because the heat dissipates from the optic itself, not the cage.
QSFP-DD arguments
- Backward compatible with QSFP/QSFP28/QSFP56 — older 100G/200G cages take QSFP-DD optics.
- More common in Ethernet-only deployments historically.
- Slightly cheaper optics on average at 400G generation.
Mixed environments
Standard pattern in 2025+ AI deployments:
- IB switches: OSFP cages.
- HCAs (NIC side): QSFP112 (NDR) or OSFP (XDR).
- Storage Ethernet: QSFP-DD or QSFP112.
- Cables: typically OSFP-to-QSFP112 (NDR) or OSFP-to-OSFP (XDR), pre-terminated.
Standardize the supply chain: pick one cable model per length, one transceiver model per reach, and order in volume. Field swaps between vendors during operations are how you get firmware-incompatibility surprises.
Splitter cables: how 800G ports become two HCAs
NDR-class IB switches have 800G OSFP ports. NDR HCAs have 400G QSFP112 ports. The standard topology uses splitter cables: one switch port lights up two HCAs.
Switch port (800G OSFP) ---- splitter -----> 400G QSFP112 (HCA on node A)
\---> 400G QSFP112 (HCA on node B)
Implications:
- Per switch port = 2 HCA links. A 64-port leaf serves 128 HCAs.
- Two failure domains share an optic at the switch end. If the OSFP optic fails, both downstream HCAs lose link.
- Splitter cables are pre-terminated in OSFP-to-2xQSFP112 form, in lengths up to ~30 m for AOC.
For rail-aligned topologies, you must verify that the two HCAs on a splitter belong to nodes that should share a leaf in the rail design. Otherwise you've inadvertently made a leaf serve two rails on the same node, which is the opposite of rail alignment.
Spare strategy
A 1000-cable fabric will see ~1-2% cable replacement per year between optic failures, AOC bend damage, and accidental disconnects. Plan accordingly.
Site spare inventory rules:
| Item | Spare ratio | Notes |
|---|---|---|
| AOC cables | 5% per length/type | Most-used lengths get larger ratio |
| DAC cables | 3% | Cheap, easy to keep on hand |
| Transceivers | 5-10% per model | Expensive but fast-replace |
| Pre-terminated trunks | 1 per length | Long lead time on replacement |
| Patch cables (LC, MPO) | 10% | Cheap, frequently lost |
| Cleaning kit refills | always | Cheap, get used up |
For each item, document:
- Where it lives (location code + bin number).
- Vendor part number for reorder.
- Lead time for fresh stock.
- Reorder threshold.
A common mistake: site goes live with full spares, two years later the spares cabinet is empty because nobody tracked depletion, and the next failure means a 6-week wait for a $50 cable. Quarterly spares audits are ten minutes; preventing one outage pays for years of audits.
Network bring-up tests by phase
Cabling commissioning happens in stages, each with its own validation.
Stage 1: physical / link-up
- All cables seated.
- Each port shows expected rate ('200G' for HDR HCAs, '400G' for NDR, etc).
- Link width is full lanes (not 'x1' or 'x2' which would indicate a bad lane).
- No optical alarms (Tx/Rx within window).
Tools: ibstatus, ibstat, ethtool, switch CLI.
Stage 2: counters clean
- Run for 10-30 minutes idle.
- Per-port FEC corrections: low and stable (not climbing).
- Symbol errors, link downs: zero.
- No port flapping.
Tools: ibdiagnet, switch counter commands, mlxlink -p for port-level details.
Stage 3: bandwidth
ib_send_bw,ib_write_bwbetween adjacent nodes — should hit 90-95% of line rate.- See perftest validation for the full test plan.
Stage 4: NCCL
- Single-node NCCL ring test (see NCCL tests).
- 2-node NCCL allreduce.
- Full-cluster NCCL allreduce at the published reference bandwidth.
If stage 1 has ports not coming up, fix them before stage 2. If stage 2 has rising errors, fix them before stage 3. Skipping ahead means debugging at scale instead of debugging at port-pair level.
Cabling SOPs that prevent the most common incidents
A short list of standard operating procedures every site should have written down.
SOP-1: Cable label format
[YYYY-MM]-[FROM_LOCATION]-[FROM_PORT]-[TO_LOCATION]-[TO_PORT]
e.g. 2025-04-rk04n07-hca0-rk09leaf01-p34
Format includes install month so you can spot cables installed during a specific batch (useful when a bad cable batch is identified).
SOP-2: Cable installation order
For each cable:
- Print labels at both ends.
- Inspect connectors, clean if needed.
- Route through trays with strain relief.
- Seat at port-A end first (the "fixed" end, e.g. switch).
- Seat at port-B end (the "node" end), with service loop.
- Affix labels.
- Verify link in switch CLI within 60 s of full seating.
- Photograph the seated cable bundle.
Skipping step 7 is how you find dead ports two days later.
SOP-3: Cable removal
For each cable being removed:
- Identify by label, verify in CMDB.
- Photograph current state.
- Disconnect at node end first.
- Disconnect at switch end.
- Remove from cable tray with bundle re-banding.
- Update CMDB to mark cable as removed.
- If reusing, inspect for damage; if scrap, dispose.
SOP-4: Adding a cable to an existing rack
- Verify destination ports are not already in use (CMDB and physical).
- Pre-print labels.
- Order cable to length if not in spares.
- Schedule a maintenance window (even for "non-disruptive" adds — there's always a small risk of disturbing existing cables).
- Install per SOP-2.
- Update CMDB and rack elevation diagram.
The discipline pays off in years.
Common physical-network mistakes
- Mixing OSFP and QSFP-DD optics in the same cage. They're not compatible; the optic doesn't seat or doesn't link. Verify cage/optic mating before bulk ordering.
- Buying MMF when you should have bought SMF. MMF saved money on day one but you can't go past 100 m and you can't run higher speeds. SMF is future-proof; the optic cost difference shrinks at higher speeds.
- Under-counting the leaf-to-spine cables. Easy to forget when sizing the BOM. They're as numerous as server-to-leaf cables in a full-bisection design.
- AOCs with too-tight bend radius. AOC is a permanently terminated cable; you can't replace just the optic. A kinked AOC is a thrown-away cable.
- Polished fiber connector (PC) mating with angled (APC) connector. Different polish, different reflection profile. Connectors look similar; APC is green-keyed, PC is blue. Match colors before mating.
- One tray for IB and Ethernet. Mixed bundles are hard to trace. Separate trays from day one.
- Color-blind operators on a color-coded rail policy. Use distinct colors and use printed text labels too. Don't rely on color alone.
- No spares on site. First cable failure ends in a 2-day outage waiting for a $50 cable. Keep 2-5% of every cable type as on-site spare.
Pre-terminated trunks vs field-terminated cables
For longer fiber runs (5-50 m or more) you have a choice:
Pre-terminated trunk
A factory-built fiber bundle with MPO connectors at each end, typically 12 or 24 fibers, ordered to a specific length. You pull it through the building, plug both ends into MPO patch panels, done.
Pros:
- Factory polish quality is consistent.
- Installation time: roughly 1 trunk per hour (pull, label, test).
- No on-site polishing equipment needed.
Cons:
- Length must be specified accurately (you can't easily shorten a pre-terminated trunk; you can only order shorter).
- Lead time: 2-6 weeks for custom lengths from quality vendors.
- More expensive per fiber-meter than spool-and-terminate.
Field-terminated
A spool of bare fiber + on-site termination kit. You measure, cut, polish, and connectorize on the floor.
Pros:
- Custom length per cable, no waste.
- Cheaper per meter for high volume.
- Can repair a damaged connector in place.
Cons:
- Termination quality depends on the technician. A sloppy hand-polish gives 1+ dB excess loss.
- Slow: 15-30 min per connector pair for fusion splicing or epoxy-polish.
- Requires polishing kit, fusion splicer, inspection scope on site.
For AI builds, pre-terminated trunks are the standard for backbone runs (row-to-row and aisle-to-aisle), and AOCs are standard for in-row links. Field-terminated is reserved for one-off custom situations.
Cable plant documentation
A cable plant without documentation is a cable plant you can't operate. The minimum documentation set:
Cable database
Per cable:
- Unique cable ID (e.g.,
2025-04-rk04n07-hca0-rk09leaf01-p34). - Source endpoint (rack, node, port).
- Destination endpoint (rack, switch, port).
- Cable type, length, speed, color.
- Install date.
- Last verified date.
- Optic part numbers (for fiber + transceiver links).
For 5000 cables this is a relational table with ~10 columns. Maintain it. Audit it quarterly.
Patch panel diagrams
Per panel:
- Panel location (rack, U-position).
- Port-by-port assignment: which port maps to which trunk fiber.
- Connector type (LC, MPO).
- Polish (UPC vs APC).
Patch panels are where cable plants get confusing. A well-documented panel is the difference between "connect to port 14" and "find the right cable in this mess."
Topology drawings
A logical drawing showing the fabric: every node, every switch, every link. Updated as the cluster grows. Required reading for any engineer who works on the network.
For small clusters (single SU), one diagram is enough. For larger ones, layered diagrams: pod-level, row-level, rack-level.
Hand-off and acceptance
When the cabling vendor finishes and hands off the build, the acceptance package should include:
- Cable database export (CSV / Excel / DCIM import-friendly format).
- Patch panel diagrams (PDF).
- Photographs of every rack, front and rear.
- Test results: link state per port, optical power per fiber link, BER results.
- BOM with serial numbers (or at minimum batch / lot numbers) for transceivers.
- Spare inventory location and counts.
If any of these are missing, do not sign off. The hand-off is the only time the cabling vendor will give you this much information. Three months later, when something goes wrong, you'll need it.
Aging cables: when do you replace?
Optical cables and transceivers don't last forever, and aging is a slow process you can plan for instead of react to.
Lifetime expectations
| Component | Typical service life | Wear-out signal |
|---|---|---|
| Passive DAC | 7-10 years | rare; usually mechanical (pulled apart) |
| AOC cable | 5-7 years | rising FEC corrections, optical power drift |
| Transceiver (modern PAM4) | 5-7 years | gradual Tx power drop, rising RX errors |
| Fiber patch (SMF, properly cared for) | 15-20+ years | only with mechanical damage |
| MPO trunk | 15-20+ years | connector wear from repeated mating |
| Cassettes / patch panels | 10-15 years | mechanical wear |
Proactive replacement strategy
For optics: track a Tx power trend per port. When Tx drifts to within 1 dB of the spec floor, replace at next maintenance window — don't wait for failure.
For AOCs: track FEC correction rate per link. A link with rising corrections (over weeks) is a candidate for replacement.
For long-haul fiber: there's no proactive trigger short of insertion-loss measurement — but insertion loss can drift due to connector contamination accumulating, so quarterly cleaning and remeasure catches it.
Replacement that doesn't disrupt jobs
For rail-aligned fabric, a single cable replacement affects one rail of one node. NCCL can route around it for some collectives but not all. Standard procedure:
- Drain the affected node from the workload (drain in Slurm, cordon in K8s).
- Wait for jobs to migrate.
- Replace the cable.
- Verify link state, run a quick perftest pair test.
- Return node to scheduling.
Total impact: ~30-60 min per cable, all of it on one node. At 1000 cables and 1% annual replacement, that's ~10 nodes per year, totally tolerable.
Inventory database integration
Modern AI sites integrate cabling into the broader DCIM:
- Cable database in DCIM (Sunbird, NetBox, custom).
- Topology synced from cable database to monitoring (so you can correlate "leaf-04 port 32 had FEC errors" to "node-rk04n07 hca0").
- Spare inventory linked to cable database (when a cable is replaced, the database tracks which spare was consumed).
- Photographs linked to rack records.
- Maintenance log tied to cable IDs.
A site running this integration finds a failing port in minutes and replaces it in under an hour. A site without it finds the port in hours and is unsure what spare to grab.
See also
- Rack design — where these cables physically route
- Power — power side of the same cable plant decisions
- Cooling — thermal constraints affecting cable trays and optics
- InfiniBand architecture — fabric topology this cable plant builds
- InfiniBand implementation — bring-up sequence using these cables
- PCIe topology — node-internal layout that determines which HCA goes to which leaf
- Health check runbook — daily checks that include link state and optical health
Sources
- NVIDIA Cable Management Guidelines and FAQ: cable types, transceiver power, reach numbers per speed.
- NVIDIA 400G (100G-PAM4) OSFP & QSFP112-based Cables and Transceivers User Guide: NDR-era reach and form factor specifications.
- NVIDIA DGX SuperPOD H100 reference architecture: 32-node SU cable counts, leaf/spine port budgets.
- IEEE 802.3 / IEC standards for 400GBASE and 800GBASE optical reach.
- Open Compute Project ORv3 specifications: in-rack busbar, blind-mate compute trays.
- ASHRAE TC 9.9: thermal constraints affecting transceiver lifetime in hot-aisle racks.