When the Network Became the Storage Bottleneck

In May we added a pain-point node called the Data Loading Bottleneck: AI training and inference sitting GPU-idle while object storage delivers the next batch — up to ~80% of end-to-end training time lost to storage waits at hyperscaler workloads. In June we wrote the Training I/O Tax about the storage tier being repriced by the GPU. The node predates everything in this post, which is the point of keeping an ontology: when the industry finally moved, we had the map square it lands on.

Here is the move. The 2026 answer to the data-loading bottleneck is not, primarily, faster storage software. It's a rebuild of the network underneath the storage — three rebuilds, actually, racing each other — and the object-storage tier's throughput ceiling is now set by a fabric decision most storage buyers have never had to make before.

Why topology decides everything

The physics were already on this index. The TopKV topology paper measured the ladder: NVLink moves ~900 GB/s inside a node, InfiniBand ~50 GB/s across nodes, cross-datacenter TCP collapses to ~12.5 GB/s — a 72× spread ([independent-methodology] academic measurement).1 Its corrected v2 figure puts a 70B model's KV cache at 1.3 GB per request at 4K context. The September verification pull confirms both numbers still stand: no v3 revision, no rebuttal — as of early September this remains the most accurate public model for KV-cache data movement.1 A 1.3 GB object that has to cross the wrong rung of that ladder stops being a cache and starts being a liability.

So the fabric is not plumbing. For the workloads this index tracks — training reads, checkpoint storms, KV-cache transfer between disaggregated prefill and decode pools — the fabric is the storage product's performance envelope.

Three futures for the wire

Ultra Ethernet. The UEC 1.0 specification shipped June 11, 2025 ([vendor-published] consortium release), and its Ultra Ethernet Transport is a genuine redesign rather than RoCEv2 with better tuning: scalable sender-based congestion control, plus multi-path packet spraying — data fragmented across every available network path simultaneously and reordered at the destination.2 That kills the "elephant flow" collision on a single switch port, and with it the incast-congestion collapse that throttles exactly the traffic pattern object storage generates: massive parallel reads. This is no longer slideware — Broadcom's Tomahawk 6 ASIC is shipping with 102.4 Tbps of switching capacity built for these features ([vendor-published] hardware specification).2

InfiniBand XDR. NVIDIA's Quantum-X800 platform — Q3400-RA switch plus ConnectX-8 SuperNICs — doubles per-port bandwidth over NDR to 800 Gbps, with 144 ports per switch and a 9× jump in in-network computing (14.4 Tflops via SHARPv4) ([vendor-published] hardware metrics).3 The deployments are real: CoreWeave brought the first NVIDIA Vera Rubin NVL72 rack online June 1, 2026, connected over Quantum-X800 ([vendor-published] deployment data), and Microsoft Azure is deploying the same fabric under its GB200 clusters.4

RoCEv2, eating the middle. Between the two, Ethernet-with-RDMA keeps taking share on cost, supply-chain diversity, and the absence of proprietary lock-in. Dell'Oro Group projects Ethernet surpasses InfiniBand in AI back-end networks by 2027 ([independent-methodology] market forecast).2 NVIDIA is hedging both sides of its own bet — the Spectrum-X800 Ethernet platform ships RoCEv2 alongside InfiniBand at CoreWeave-class facilities, at up to 1.6 Tb/s of backend bandwidth per GPU ([vendor-published]).3

A provenance note, because that's the house style: the widely repeated claim that "RoCEv2 reaches 85–95% of InfiniBand throughput at lower cost" — a figure the RDMA node on this index carries — came back from the primary-source sweep with no independent measurement attributable. The premise is validated by vendor behavior; the specific percentage traces to tertiary comparison content. The node now says so. Treat it as directional, not measured.

The storage angle: audited numbers and an upstream merge

What do these fabrics buy the bucket? Two data points from this wave, one audited and one architectural.

The audited one: in MLPerf Storage v2.0 results, DDN's AI400X3 sustained 120.68 GB/s on 3D U-Net training workloads and processed Llama3-8b checkpoints at 30.6 GB/s read / 15.3 GB/s write — from a 2RU appliance ([independent-methodology] benchmark, MLPerf-audited).5 The contrast is as informative as the number: WEKA did not submit audited figures for the v2.0 round, leaving its 10.2 TB/s rack-scale claim unverified by independent auditors.5 In a market where every vendor slide says "GPUDirect-ready," audited versus vendor-published is becoming the primary axis of comparison — the same lesson blog #29 drew about arXiv preprints, replayed at the appliance tier.

The architectural one: Dell integrated GPUDirect RDMA KV-cache offloading into ObjectScale, its S3-compatible platform — and the capability is merged upstream into both vLLM and LMCache, supporting S3-HTTP and S3-over-RDMA paths ([vendor-published] release notes).6 Read that against the ledger: our August 24 call said object storage is reclassifying into AI's active memory tier, and our August 28 call wanted a second named production deployment before the Cohere/CoreWeave case counted as a pattern. Dell ObjectScale is that second named platform. What Dell hasn't published yet is latency numbers, which is what the call's clause literally asks for — so the calls ledger records it as progress, not satisfaction. Calls get upgraded by evidence, never silently.

The economics forcing the choice

By late 2025, networking equipment cost and power consumption began to rival the GPUs themselves ([independent-methodology] financial analysis).7 That drags the fabric decision up to where storage TCO already lives. CoreWeave claims its Vera Rubin deployment with optimal networking delivers up to 10× lower cost per million tokens for agentic inference versus earlier Blackwell deployments ([vendor-published] claim).4 Moving to UEC-class Ethernet cuts the InfiniBand per-port premium, and the Intersect360 analysis makes the reallocation explicit: capital freed from the switch line goes into denser high-performance object storage — without giving up the RDMA path GPUDirect requires.2

Two adjacent pressures keep this from being optional. The NAND supply crisis has broken into a plateau — TrendForce's Q3 survey shows contract prices moderating to +10–15% QoQ after the +55–60% and +70–75% quarters ([independent-methodology] market forecast)8 — but a plateau at a permanently elevated baseline, which locks in the tier-to-object-storage architecture rather than relieving it. (One reconciliation worth printing: Q2's $37.59B at +103.6% QoQ is the top five enterprise SSD vendors' segment revenue; the $68.87B at +77% QoQ we cited in August is the top five NAND vendors' total revenue. Two true numbers, two different metrics — the node now keeps them apart.8) And datacenter power is physically relocating compute: CoreWeave's Stockholm deployment with Conapto runs on 100% renewables with heat recovery into the city's district heating; Nscale built in Stavanger ([vendor-published] press releases).9 When the GPUs move to power-rich regions, the training data and KV caches either migrate with them or cross exactly these fabrics between zones — congestion management as a storage-procurement line item.

The ledger keeps us honest

The same research pull verified the August posts' claims, and one expectation didn't survive: DuckDB 2.0 has not shipped. It remains in Preview ("A Preview of DuckDB v2.0," August 17); current stable is v1.5.5 (July), LTS v1.4.5 (June) ([vendor-published] release notes).10 Our published node said "tracked for Fall 2026," which stands — but our September research brief assumed a September ship, and blog #29's "Quack protocol going stable" phrasing read nearer-term than reality. The correction is logged on /calls, where our misses live in public.

The lakehouse front — the wave's second research thread — folded into nodes rather than carrying the post: the SAP acquisition of Dremio closed July 6, 2026, with Dremio's Apache Iceberg, Polaris, and Arrow commitments maintained, as one leg of a Dremio+Reltio+Prior Labs structured-data pipeline ([vendor-published] press releases); the Iceberg V3 node gained the documented v2→v3 migration gotcha (metadata-only upgrade, but upgrade every engine first — a partial write from a v2-only engine can corrupt table state); DataFusion Comet hit 1.0.0 after two years of incubation; and Data Contracts now points exclusively at ODCS v3.1.0, the competing DCS having formally deprecated itself.

Our read

Dated call, 2026-09-03: by end of 2027, the fabric choice becomes an object-storage product feature — Ethernet-family fabrics (UEC/RoCEv2) carry the majority of new S3-to-GPU deployments, and S3-compatible vendors ship fabric-specific tuning as named product capabilities rather than reference-architecture PDFs. We revise if UEC silicon fails to reach named production storage deployments by mid-2027, if InfiniBand breaks the Dell'Oro crossover by holding share through 2027, or if S3-over-RDMA stays a two-vendor story without a second round of audited results. Logged on the calls ledger.

What changed on the index

Five new nodes: Ultra Ethernet (UEC) and NVIDIA Quantum-X800 — the index's first network-fabric nodes, closing a gap our own orphan-signal pipeline flagged; Dell ObjectScale, the second named S3-tier KV-cache platform; DDN Infinia, carrying the MLPerf-audited throughput reference points; and Open Semantic Interchange (OSI), the semantic-model interchange standard (source-noted — launch coverage is still single-source). Enrichments landed on DuckDB (2.0 not-shipped status), NAND Flash Supply Crisis (Q3 plateau + metric reconciliation), MCP (Stacklok's 66% RBAC / 51% audit-logging survey alongside the confirmed static-key figures), LMCache and NVIDIA GPUDirect RDMA for S3 (the ObjectScale corroboration), RDMA (UEC status + the 85–95% provenance note), Dremio, Iceberg V3 Spec, DataFusion, Data Contracts, Backblaze B2 ($6.95/TB with API fees eliminated — the load-bearing change for small-object AI workloads), CoreWeave AI Object Storage, Data Loading Bottleneck, LiteLLM (final TeamPCP scope: 1,000+ enterprise environments), and Datacenter Power Shortfall. One correction and two call updates on /calls. The index stands at 437 nodes.

Works cited

Footnotes

  1. Topology-Aware Data Movement for Disaggregated GPU Inference, v2 (arXiv 2607.28633) — the bandwidth ladder and the corrected 1.3 GB/request figure; no v3 or rebuttal as of early September 2026. 2

  2. Ultra Ethernet Consortium — Specification 1.0 launch and Intersect360 UEC 1.0 white paper; Tomahawk 6 shipping status and the Dell'Oro forecast per Nokia — the future of AI networking with UEC. 2 3 4

  3. NVIDIA — new switches for trillion-parameter GPU computing and Introl — Quantum-X800 and the XDR generation. 2

  4. CoreWeave — industry-first bring-up of NVIDIA Vera Rubin NVL72; Azure GB200 deployments per Microsoft Azure — NVIDIA partnership. 2

  5. StorageReview — best storage arrays 2026: AI leaders and audited results — MLPerf Storage v2.0 audited DDN numbers and the WEKA non-submission. 2

  6. Dell — KV cache offload to object storage: GPU-Direct and upstream — the vLLM/LMCache upstream merge.

  7. MarketMinute — the rewiring of Silicon Valley: trading GPUs for interconnects.

  8. TrendForce — AI server demand continues to support memory prices in 3Q26 and TrendForce press center — Q2 2026 revenue figures; LTA dynamics per TrendForce — long-term agreements cap price increases. 2

  9. CoreWeave — Conapto partnership expands AI cloud capacity in Sweden; Nscale Stavanger per Nokia — UEC.

  10. DuckDB releases (GitHub) and DuckDB engineering blog — v1.5.5 stable, v1.4.5 LTS, 2.0 in Preview.