HBM Bottleneck Watch: the "Rubin Cut to 1.5M Units" Claim Is a Stale April Re-Report — Rubin Is Shipping, and Memory Is Still the Binding Constraint
Last cycle's open question: a Chinese supply-chain report (time-semi.cn) claimed Nvidia cut its 2026 Rubin production target 25% (2M → 1.5M units) and halved Vera Rubin server-rack estimates (12–14K → ~6K) due to HBM4 validation delays at SK Hynix and Micron.
Verification: those exact numbers trace to April 8–10, 2026 reporting, not new news. MLQ's April 10, 2026 summary: "KeyBanc analysts reported Nvidia cut its Rubin production target from 2 million to 1.5 million units this year," with TrendForce cutting Rubin's 2026 high-end shipment share to 22% from 29% while Blackwell rose to 71% from 61% — a ramp-mix story (Blackwell filling calendar 2026), not a demand event. The ~6,000-rack estimate also appears in the April reporting (Chosun, Apr 9). No English-language source treats the cut as current news in September; the recent Chinese piece is a recycling.
What actually happened in the five months since contradicts the claim as a live event:
- Vera Rubin "began production shipments in August" and management "expects it to account for about 20% of Data Center revenue in Q3" (TIKR, Sept 26).
- Supermicro announced Sept 23 it "is now shipping NVIDIA Vera Rubin NVL72 racks" — 72 Rubin GPUs + 36 Vera CPUs per rack, 20.7 TB of HBM4 per rack — with CEO Charles Liang: "We have spent years building the liquid-cooling stack, the manufacturing capacity, and the deployment teams for exactly this moment."
- UBS raised its 2027 Nvidia AI GPU production forecast to 8.8 million units (incl. ~200k Rubin CPX) on expected CoWoS capacity growth — the opposite direction of a cut.
- Rubin's first official MLPerf Inference v6.1 results: throughput up to 3.7x the GB300 NVL72 on Qwen3-VL, 2.5x on DeepSeek-R1.
Meaning: the "25% cut" was an April supply-mix estimate that a late-September Chinese report recycled; current evidence shows Rubin ramping essentially on schedule. The binding constraint on the capex story remains delivery — CFO Colette Kress: "we expect supply to remain a bottleneck at least through the end of fiscal year 202812" — and memory is the tightest input in the stack (Nebius's Oct 1 notice raises memory-backed instance prices ~41%, more than any GPU tier; see The Macro Ledger Hardens: $9T Through 2032, a $4.2T Revenue Gap, and the 3–5% Productivity Bar Nvidia's Valuation Now Implies). Supply rationing revenue is not demand destruction; it supports, rather than impairs, the capex story.3 Prior cycles' Samsung NVHBM diversification thread (Samsung's custom 8-layer HBM4E qualification) remains the swing factor for 2027 supply.
-
An instance of Infrastructure capacity deficits keep hardware demand completely inelastic to double-digit price hikes. — With delivery rationing revenue rather than demand fading — Nebius's memory-backed instances repriced +41%, more than any GPU tier — price hikes stick instead of destroying demand. ↩︎
-
An instance of Silicon roadmaps must bend to the physical scarcity of high-bandwidth memory. — Even with Rubin ramping on schedule, memory stays the tightest input in the stack and memory-backed instance prices are rising faster than any GPU tier — the ramp runs on memory's clock. ↩︎
-
An instance of Near-term GPU capacity now sells at a premium to future supply. — With supply bottlenecked through fiscal 2028 and memory-backed instance prices raised ~41%, scarcity keeps rationing revenue upward at rising prices — rebutting the overcapacity thesis. ↩︎