//
Podcast · 2026-09-14

SemiAnalysis Weekly: Ep. 030 - Long Live the Short King: Why 4-HI HBM Wins (Memory) | Myron Xie, Jordan Nanos

Rubin Ultra cut from 1TB to 192GB HBM: why supply — not performance — says 4-high wins.

Episode

SemiAnalysis Weekly
ShowSemiAnalysis Weekly
EpisodeEp. 030 - Long Live the Short King: Why 4-HI HBM Wins (Memory) | Myron Xie, Jordan Nanos
GuestMyron Xie
HostJordan Nanos
Published2026-09-14T17:00:16Z
Duration39 min
Fidelity[partial]
StanceNEUTRAL
ListenEpisode link

Abstract

Myron Xie, SemiAnalysis analyst, and host Jordan Nanos explain why NVIDIA’s Rubin Ultra will ship with 192GB of HBM4 per package instead of the previewed 1TB: a supply-driven cut from 12-high to 8-high HBM cubes — and from four compute dies to two. They argue 4-high HBM is the rational endpoint, since bandwidth is fixed per stack while memory vendors price per gigabyte, and inference workloads are bandwidth-bound with diminishing returns to extra capacity. The discussion traces the cascade into DRAM, logic, substrate and power bottlenecks, why memory supply cannot expand this decade, and the counter-case that exploding model sizes would flip the economics back to taller stacks.

The Theses

8 claims
1. Rubin Ultra’s 1TB → 192GB HBM cut is supply rationing, not a performance choice.
Evidence: Myron: the HBM secured would be insufficient, shipped as 12-high cubes, to cover all the logic secured at TSMC — so NVIDIA rations DRAM supply into 8-high cubes to maximize accelerator shipments.
2. Four-high is the physical and economic floor for HBM stacks.
Evidence: At least four DRAM layers are needed for a cube’s full bandwidth (bandwidth is fixed per stack, driven by the interface), while suppliers price per gigabyte — so dollar-per-bandwidth is far better at 4-high than 8- or 12-high.
3. Inference has displaced pre-training as the dominant share of frontier-lab compute, tilting systems toward bandwidth over capacity.
Evidence: SemiAnalysis’s tokonomics data-center model shows the marginal data center entering lab fleets goes to inference/post-training — “not pre-training runs for that 20 trillion parameter model that we were thinking about.”
4. Model parameter counts are not scaling as aggressively as the Rubin roadmap assumed, which underwrites the capacity cut.
Evidence: Kimi K3 sits at ~2.8T parameters (~6× Llama 3.1 405B), MXFP4 halves capacity needs, one K3 weight set is ~8% of an NVL72 GB200 domain’s HBM — and NVL576’s 8× larger scale-up domain gives ample aggregate capacity.
5. Dropping to 4-high more than doubles harvestable HBM cubes versus 8-high, shifting the bottleneck elsewhere.
Evidence: Per-layer yield loss compounds 8–12× versus 4×, and powering a 4-die stack is easier — so the constraint moves from HBM to leading-edge logic wafers (TSMC), substrates, and power.
6. Relaxing HBM frees DRAM wafers the whole server industry needs.
Evidence: Conventional DRAM is tight enough that servers are being de-specced per socket, partly because DRAM wafers are cannibalized for HBM — freeing them helps commodity DRAM that AI servers need for CPU-bound work like tool calls.
7. Memory supply will not ease “within this decade.”
Evidence: Cleanroom construction lead times and ASML’s EUV tool supply chain — specialized suppliers for mirrors and the like — are hard physical constraints on new DRAM wafer capacity.
8. Expect multiple memory-capacity SKUs of the same logic die.
Evidence: The capacity-vs-supply tradeoff varies by customer — Meta already runs a custom MI450 variant with 8-high instead of 12-high HBM.

Key Math

  • 1TB → 192GB — Rubin Ultra per-package HBM, previewed vs shipping (the first NVIDIA flagship with less capacity than its predecessor) Myron (~02:25)
  • 4 → 2 compute dies per package Myron (~03:28)
  • 16 stacks of HBM4E at 16-high → HBM4 at 8-high Myron (~03:38)
  • 288GB — Blackwell Ultra and vanilla Rubin, 12-high HBM Myron (~04:02)
  • 8 × 80GB = 640GB — Hopper HGX server HBM Myron (~08:54)
  • 45GB at FP8 — Llama 3.1 405B footprint, ~60% of an HGX server’s HBM Myron (~09:19)
  • ~2.8T parameters — Kimi K3, ~6× the size of Llama 3.1 405B; MXFP4 quantization halves capacity needs Myron (~09:53)
  • 72 GPUs × 288GB — NVL72 GB200; one K3 weight set is ~8% of the domain’s HBM Myron (~10:32)
  • NVL576 — Rubin Ultra scale-up domain, ~8× larger than NVL72 Myron (~12:14)
  • ~2,048 data I/Os per HBM4 stack between cube and compute; ≥4 DRAM layers needed to use the full interface (auto-caption garbled the figure) Myron (~17:25)
  • 99% — illustrative per-layer yield; compounding ~1% loss 8–12× vs 4× is why 4-high yields more than double the harvestable cubes of 8-high Myron (~23:12)
  • 3× K3 — modeled model size at which 8/12-high’s extra capacity becomes worth the cost Myron (~33:33)
  • V100 16→32GB; A100 40→80GB; H100 80GB → H200 144GB — prior generations got capacity-bump SKUs, precedent for multiple SKUs Jordan (~36:12)
  • 10T+ parameter models — the talk when Rubin was announced; looped transformers (confirmed in GPT-6 Astra) add compute depth without parameters Jordan (~07:25)

Quotes

Reuben Ultra was going to be a terabyte and it's now being revised to 192 gigs, which is 192 GB is less than the 288 GB that's shipping today that people are using today.

Jordan Nanos  ·  SemiAnalysis Weekly  ·  ~00:19:22

Why it matters: The headline revision, stated with the comparison that makes it historic.

So, speeding up tokens is uh it's all from bandwidth basically. Um the capacity doesn't really uh add anything.

Myron Xie  ·  SemiAnalysis Weekly  ·  ~00:13:10

Why it matters: The bandwidth-over-capacity thesis in one breath — inference speed comes from the interface, not the stack height.

the marginal data center that they bring into their fleet is not going towards pre-training. It's going towards research, post-training, inference, some collection of other that is uh not pre-training runs for that 20 trillion parameter model that we were thinking about.

Myron Xie  ·  SemiAnalysis Weekly  ·  ~00:15:26

Why it matters: The demand-side evidence that capacity-hungry pre-training no longer drives incremental HBM demand.

I mean we don't uh I think we subscribe to the memory model to find more but um I'd say the TLDDR is not uh not within this decade

Myron Xie  ·  SemiAnalysis Weekly  ·  ~00:31:55

Why it matters: SemiAnalysis’s own timeline for memory-supply relief, stated flatly (auto-caption artifact verbatim).

it would be very embarrassing for OpenAI and Anthropic if Kim K3 was this close to performance with them.

Jordan Nanos  ·  SemiAnalysis Weekly  ·  ~00:35:20

Why it matters: The tell on frontier model sizes — implied parity with an open ~2.8T-param model undercuts the bigger-model thesis.

Variant Perception

Priced in: HBM tightness and rising memory prices are consensus; so is NVIDIA’s design leverage over suppliers. The market already treats HBM supply — not logic — as the gating item for AI accelerator shipments.

What’s new: the magnitude of the Rubin Ultra revision (1TB → 192GB, four compute dies to two, HBM4E to HBM4) and the mechanism — NVIDIA is actively rationing DRAM into 8-high cubes to stretch HBM across more logic, with 4-high as the endgame. The “tokens per HBM wafer” framing makes the vendor math explicit: maximize aggregate tokens, not per-chip capacity.

The bear case: SemiAnalysis’s own counter — if model sizes explode (their 3× K3 model), the batching economics flip back to favoring taller stacks, and a 4-high fleet becomes suboptimal. Labs’ “loudest cries” are for bandwidth today, but roadmaps change; and falling per-chip HBM content is a headwind for memory vendors’ revenue per accelerator.

Discount: SemiAnalysis sells to industry readers and kept the memory-supplier impact “behind the paywall” — the analysis is deepest where it is most promotable. The “not within this decade” timeline is their house view, stated without the underlying model shown; and auto-caption artifacts mean several technical figures are approximate.

Why It Matters

This reframes the AI hardware bottleneck from “how fast can NVIDIA design” to “how fast can DRAM wafers grow” — and the answer is they can’t, this decade. If vendors optimize for tokens per HBM wafer instead of per-chip capacity, the winners are whoever ships the most logic through a fixed HBM allocation (TSMC leading-edge wafers, substrate and power become the next constraints), while memory vendors sell more cubes at lower dollars-per-cube. The open question the episode leaves is whether frontier labs’ models stay small enough for 4-high to hold — if they don’t, today’s “rational endpoint” becomes tomorrow’s regret.

Positioning Read

Directional only
SUPPORTED
AI capex supercycle
Supply-driven rationing — cutting HBM per chip to ship more accelerators through a fixed DRAM allocation — is demand-outrunning-supply in physical form; the episode’s whole premise is that demand exceeds what the memory supply chain can feed.
SUPPORTED
Power & interconnect is the binding constraint
When 4-high more than doubles harvestable cubes, the binding constraint moves from HBM to leading-edge logic wafers, substrates, and power — the constraint stack deepens rather than clearing.
SUPPORTED
Inference economics favor consumption pricing
Inference and post-training dominating frontier compute share, plus tokens-per-HBM-wafer optimization, is the physical counterpart of consumption-priced inference economics.

Frameworks

Tokens per HBM wafer
In a DRAM-constrained world, optimize the scarce resource the way the industry optimizes tokens per watt and tokens per dollar: 4-high HBM is the configuration that maximizes aggregate tokens from a fixed HBM-wafer allocation. (SemiAnalysis Weekly Ep. 030, 2026-09-14)