Rubin Ultra cut from 1TB to 192GB HBM: why supply — not performance — says 4-high wins.
| Show | SemiAnalysis Weekly |
| Episode | Ep. 030 - Long Live the Short King: Why 4-HI HBM Wins (Memory) | Myron Xie, Jordan Nanos |
| Guest | Myron Xie |
| Host | Jordan Nanos |
| Published | 2026-09-14T17:00:16Z |
| Duration | 39 min |
| Fidelity | [partial] |
| Stance | NEUTRAL |
| Listen | Episode link |
Myron Xie, SemiAnalysis analyst, and host Jordan Nanos explain why NVIDIA’s Rubin Ultra will ship with 192GB of HBM4 per package instead of the previewed 1TB: a supply-driven cut from 12-high to 8-high HBM cubes — and from four compute dies to two. They argue 4-high HBM is the rational endpoint, since bandwidth is fixed per stack while memory vendors price per gigabyte, and inference workloads are bandwidth-bound with diminishing returns to extra capacity. The discussion traces the cascade into DRAM, logic, substrate and power bottlenecks, why memory supply cannot expand this decade, and the counter-case that exploding model sizes would flip the economics back to taller stacks.
Reuben Ultra was going to be a terabyte and it's now being revised to 192 gigs, which is 192 GB is less than the 288 GB that's shipping today that people are using today.
Why it matters: The headline revision, stated with the comparison that makes it historic.
So, speeding up tokens is uh it's all from bandwidth basically. Um the capacity doesn't really uh add anything.
Why it matters: The bandwidth-over-capacity thesis in one breath — inference speed comes from the interface, not the stack height.
the marginal data center that they bring into their fleet is not going towards pre-training. It's going towards research, post-training, inference, some collection of other that is uh not pre-training runs for that 20 trillion parameter model that we were thinking about.
Why it matters: The demand-side evidence that capacity-hungry pre-training no longer drives incremental HBM demand.
I mean we don't uh I think we subscribe to the memory model to find more but um I'd say the TLDDR is not uh not within this decade
Why it matters: SemiAnalysis’s own timeline for memory-supply relief, stated flatly (auto-caption artifact verbatim).
it would be very embarrassing for OpenAI and Anthropic if Kim K3 was this close to performance with them.
Why it matters: The tell on frontier model sizes — implied parity with an open ~2.8T-param model undercuts the bigger-model thesis.
Priced in: HBM tightness and rising memory prices are consensus; so is NVIDIA’s design leverage over suppliers. The market already treats HBM supply — not logic — as the gating item for AI accelerator shipments.
What’s new: the magnitude of the Rubin Ultra revision (1TB → 192GB, four compute dies to two, HBM4E to HBM4) and the mechanism — NVIDIA is actively rationing DRAM into 8-high cubes to stretch HBM across more logic, with 4-high as the endgame. The “tokens per HBM wafer” framing makes the vendor math explicit: maximize aggregate tokens, not per-chip capacity.
The bear case: SemiAnalysis’s own counter — if model sizes explode (their 3× K3 model), the batching economics flip back to favoring taller stacks, and a 4-high fleet becomes suboptimal. Labs’ “loudest cries” are for bandwidth today, but roadmaps change; and falling per-chip HBM content is a headwind for memory vendors’ revenue per accelerator.
Discount: SemiAnalysis sells to industry readers and kept the memory-supplier impact “behind the paywall” — the analysis is deepest where it is most promotable. The “not within this decade” timeline is their house view, stated without the underlying model shown; and auto-caption artifacts mean several technical figures are approximate.
This reframes the AI hardware bottleneck from “how fast can NVIDIA design” to “how fast can DRAM wafers grow” — and the answer is they can’t, this decade. If vendors optimize for tokens per HBM wafer instead of per-chip capacity, the winners are whoever ships the most logic through a fixed HBM allocation (TSMC leading-edge wafers, substrate and power become the next constraints), while memory vendors sell more cubes at lower dollars-per-cube. The open question the episode leaves is whether frontier labs’ models stay small enough for 4-high to hold — if they don’t, today’s “rational endpoint” becomes tomorrow’s regret.