In Brief
Since 2024, Nvidia has been publishing a four-year data center roadmap in response to unprecedented competitive pressure: hyperscalers — Google, AWS, Microsoft — refresh their in-house chips on an annual cycle. Breaking from the previous biennial rhythm, Nvidia now aligns one generation per year: Blackwell Ultra (B300) in H2 2025, Vera Rubin (R100) in H2 2026, Rubin Ultra (R300) in H2 2027, and Feynman in 2028. With each leap, claimed performance doubles or triples — but HBM4 memory supply constraints, TSMC packaging capacity, and liquid cooling requirements are becoming as critical as raw performance.
Reading note: this article covers the 2025-2028 prospective roadmap. For the current Blackwell architecture (B100/B200), the CUDA software stack, and Nvidia’s present market dominance, see GPU Nvidia and CUDA.
Why Annual Cadence? The Strategic Context
Until 2022, Nvidia followed a biennial cycle: Ampere in 2020, Hopper in 2022. Blackwell, expected in 2024, already compressed this to roughly two years since Hopper. But the real break came in 2024 with a decision formalized by Jensen Huang at the GTC 2025 keynote (March 2025): one generation per year.
The rationale comes down to competitive arithmetic. Google’s TPU v7, AWS’s Trainium 2 and 3, and Microsoft’s Maia all refresh annually, targeting low-cost inference. Each iteration narrows the performance-per-dollar gap. Nvidia responds by publishing its public roadmap as a strategic signal: Nvidia systems will improve faster than anything a hyperscaler can build in-house.
Jensen Huang justified sustained investment with a specific claim: reasoning models (like o3 and its successors) require, in his words, “100× more compute” than classic generative models for inference. Commercial argument as much as technical one — there is no independent demonstration of this multiplier — but it legitimizes buying Nvidia hardware rather than less capable alternatives.
Official timeline (GTC 2025):
- H2 2025 — Blackwell Ultra (B300)
- H2 2026 — Vera Rubin (R100)
- H2 2027 — Rubin Ultra (R300)
- 2028 — Feynman
In plain terms: Nvidia publishes a four-year plan not out of industrial transparency, but as a competitive weapon. The roadmap tells hyperscalers: “investing in your own chips means missing four generations of Nvidia improvements.” It is as much a value proposition as a commercial pressure.
Sources: GTC 2025 keynote Jensen Huang (March 2025) [S3]; NextPlatform “Nvidia Draws GPU System Roadmap Out To 2028” (March 2025) [S7]; Tom’s Hardware “Inside the AI accelerator arms race” (2025).
Blackwell Ultra (B300) — H2 2025
What It Is: A Memory Iteration, Not a New Node
It is important to understand what the B300 is before detailing its specs. Blackwell Ultra is not a new architecture. The GPU die remains identical to Blackwell: same TSMC 4NP process, same dual-die GH100 reticle-limited design, 208 billion transistors. The improvement is concentrated in memory and associated compute density.
Vocabulary
- Process node: transistor fineness in nanometers (e.g., 4NP = Nvidia’s 4 nm process). A finer node = more transistors per mm², lower power consumption, more performance per watt.
- Die: the silicon chip itself, before dicing and packaging. The B200 and B300 use the same die.
- Reticle: the UV exposure template in the lithography machine. The maximum reticle area limits single-die size (~830 mm²). Beyond that, Nvidia assembles multiple dies.
Key Specs
Memory: 288 GB of HBM3e per GPU (vs 192 GB on B200), achieved by moving from 8-Hi to 12-Hi stacks. Bandwidth stays at 8 TB/s — taller stacks increase capacity, not interface speed.
Compute [roadmap vendor claim]: 15 PFLOPS FP4 dense (vs 10 PFLOPS on B200, +50%), 20 PFLOPS FP4 sparse, 5 PFLOPS FP8.
Interconnect: NVLink 5 at 1.8 TB/s bidirectional (unchanged vs B200), NVLink-C2C at 900 GB/s for CPU-GPU coherence. TDP: 1,400 W.
GB300 NVL72 system: 72 B300 GPUs interconnected, 1,100 PFLOPS FP4 inference [roadmap vendor claim] (+50% vs GB200 NVL72), 360 PFLOPS FP8 training [roadmap vendor claim].
Use Case Angle: Inference-Weighted
The B300 is explicitly positioned by Nvidia as “inference-weighted”. More memory per GPU means long-context models (hundreds of thousands of KV-cache tokens) fit entirely in a single pass. FP64 compute (scientific high-precision) is deliberately deprioritized — a clear signal that the target market is production inference, not physical simulation or academic research.
In plain terms: the B300 is a B200 with more memory and a 50% compute boost. No new node, no new architecture. It serves as a bridge while Vera Rubin (the real leap) ramps up at TSMC.
Sources: NVIDIA Technical Blog “Inside NVIDIA Blackwell Ultra” [S1]; Tom’s Hardware “Nvidia announces Blackwell Ultra B300” (2025); NextPlatform (March 2025) [S7]; WCCFTech “NVIDIA Blackwell Ultra B300 Unleashing In 2H 2025”.
Vera Rubin (R100) — H2 2026
Vera Rubin is the first genuine architectural leap since Blackwell. The name covers two distinct components: Rubin (GPU, codename R100) and Vera (CPU, codename CV100). Assembled together, they form the Vera Rubin superchip (VR100), analogous to Grace Hopper for Hopper or Grace Blackwell for Blackwell.
The Rubin R100 GPU: Node Jump and Doubled Memory Interface
Process node: TSMC N3/N3P (3 nm). This is the first full node jump since Blackwell’s 4NP — potentially +30 to 40% performance per watt before architectural optimization.
Die architecture: two reticle-limited dies (GR100) per SXM7 package, flanked by two I/O tiles that isolate NVLink SerDes, PCIe, and NVLink-C2C logic from the compute die. 336 billion transistors (+62% vs Blackwell). 224 Streaming Multiprocessors, 5th-generation Tensor Cores.
Memory: 288 GB of HBM4 per GPU (8 stacks, 12-Hi). The real breakthrough is the doubled memory interface: 2048 bits vs 1024 bits on HBM3e, at 6.4 Gb/s pin speed (JEDEC HBM4 ceiling). Result: ~20.5 TB/s bandwidth [roadmap vendor claim] — Nvidia announces “22 TB/s”, SemiAnalysis calculates ~20.5 TB/s, the 7% gap reflects different efficiency assumptions. Either way, that is 2.8× the B200 (8 TB/s → ~22 TB/s).
Compute [roadmap vendor claim]: 50 PFLOPS FP4 dense inference, 35 PFLOPS FP4 training. ×5 inference vs Blackwell, ×3.5 training vs Blackwell.
Interconnect: NVLink 6 at 3.6 TB/s bidirectional per GPU (doubled vs NVLink 5, same 224G SerDes but twice the lanes), NVLink-C2C at 1.8 TB/s CPU↔GPU, ConnectX-9 at 1.6 Tb/s for scale-out networking. Estimated TDP: ~1,800 W [roadmap vendor claim].
Status: at the CES January 2026 keynote, Jensen Huang announced that Vera Rubin chips are “in full production” and physically showed a Rubin wafer. Customer deliveries H2 2026. [S10]
In plain terms: Vera Rubin is not a B300 with more transistors. It is a complete leap: new TSMC 3 nm node, new HBM4 memory generation, doubled memory interface, doubled interconnects. Memory bandwidth — the real bottleneck for LLM inference — jumps 2.8×.
The Vera CV100 CPU: Nvidia’s First Fully Custom CPU
Nvidia used the Grace CPU (Hopper and Blackwell) on a licensed ARM Neoverse V2 design. With Vera, this is the first time Nvidia designs a CPU from scratch: ARM v9.2-A architecture, custom “Olympus” cores (in-house design, not a standard ARM reference design).
Specs: 88 Olympus cores + Spatial Multithreading (176 threads). Announced IPC +50% vs Grace, overall performance +150% [roadmap vendor claim]. L2: 2 MB/core. Unified L3: 164 MB. Memory: 1.5 TB LPDDR5X via SOCAMM modules (improved serviceability vs SO-DIMM). CPU bandwidth: 1.2 TB/s (3× Grace).
The strategic signal is clear: Nvidia is positioning itself on the CPU layer of the AI datacenter, in direct competition with Intel Xeon and AMD EPYC for “AI factories” — datacenters built end-to-end for AI.
The NVL144 System: 144 GPU Dies, 1 Rack
Configuration: VR200 NVL144 — 72 Vera Rubin packages (= 144 GPU dies + 72 Vera CPUs). Terminology note: Nvidia counts “dies,” not packages, in the system name. SemiAnalysis and NextPlatform noted this evolution as a way to display optically larger numbers (“Jensen Math”).
System performance [roadmap vendor claim]:
- 3.6 EFLOPS FP4 inference (1 exaflop = 10¹⁸ floating-point operations/second)
- 1.2 EFLOPS FP8 training
- ×3.3 vs GB300 NVL72; ×5 vs GB200 NVL72
Fabric: NVLink 6 at 260 TB/s aggregate per NVL72 rack (all-to-all 72-GPU topology), Spectrum-6 switch at 102.4 Tb/s.
Sources: NVIDIA Technical Blog “Inside the NVIDIA Vera Rubin Platform” (January 2026) [S2]; SemiAnalysis “NVIDIA GTC 2025 — Built For Reasoning, Vera Rubin, Kyber, CPO, Dynamo Inference” [S5]; SemiAnalysis “Vera Rubin – Extreme Co-Design” [S6]; Tom’s Hardware “Nvidia’s Vera Rubin platform in depth” (2026) [S8]; DataCenterDynamics (January 2026) [S10]; VideoCardz “NVIDIA Vera Rubin NVL72 Detailed” (2026) [S13].
The HBM4 Challenge: Supply Structure and Exclusions
The HBM3e → HBM4 transition is the most critical physical bottleneck in the Rubin roadmap. Understanding who supplies the memory — and why — illuminates execution risks for the roadmap.
SK Hynix 70%, Samsung 30%, Micron Excluded from Flagship
Nvidia requires HBM4 with pin speeds above 10 Gb/s. Among the three manufacturers capable of producing HBM4 — SK Hynix, Samsung, Micron — only SK Hynix and Samsung have been qualified for the flagship Rubin R100 GPU:
- SK Hynix (~70% of volumes): full production from March–April 2026. [S11, S12]
- Samsung (~30%): qualification achieved at 10 and 11 Gb/s (Q1 2026). [S11]
- Micron: excluded from flagship Rubin. Positioned on the “Rubin CPX” mid-range segment. [S11]
This concentration across two suppliers is a structural supply risk. A disruption at SK Hynix (yield issue, factory incident) would affect ~70% of Nvidia’s HBM4 supply.
HBM4 16-Hi for Rubin Ultra: An Unresolved Constraint
For Rubin Ultra (2027), Nvidia is requesting HBM4 16-Hi stacks (16 DRAM layers instead of 12) to reach 1 TB per package. These stacks are in development at all three manufacturers, with Nvidia requesting delivery by Q4 2026. As of the writing date (April 2026), no manufacturer has published validated specifications for this configuration.
In plain terms: Rubin Ultra’s timeline depends on memory technology that is not yet in production. This is an identified, unresolved risk in Nvidia’s roadmap — one to monitor for enterprise buyers planning 2027 deployments.
Sources: TrendForce “Samsung, SK Hynix Reportedly Tapped as NVIDIA Rubin HBM4 Suppliers” (March 2026) [S11]; TrendForce “SK Hynix Reportedly to Supply About Two-Thirds of NVIDIA HBM4” (January 2026) [S12].
Rubin Ultra (R300) — H2 2027
All information on Rubin Ultra comes from the GTC 2025 announcement (March 2025). No independent measurements are available. The entirety of this section is tagged [roadmap vendor claim] unless otherwise noted.
Rubin Ultra GPU: 4 Dies per Package
The architectural breakthrough of Rubin Ultra is the move from 2 to 4 GPU dies per SXM8 package. The GPU die likely remains the same type as Rubin R100 (high probability — unconfirmed), but the package contains 4 dies instead of 2, doubling the GPU silicon per unit purchased.
Compute [roadmap vendor claim]: 100 PFLOPS FP4 dense per GPU package (×2 vs Rubin R100).
Memory [roadmap vendor claim]: 1 TB of HBM4E (HBM4 Enhanced, pin speed >8 Gb/s) per GPU package, via 16 stacks 8-Hi HBM4E. Estimated bandwidth: ~40 TB/s per package.
Interconnect [roadmap vendor claim]: NVLink 7 (GTC 2025 roadmap version — detailed specs not disclosed), NVSwitch 6 at 7.2 TB/s. Estimated TDP: ~3,600 W per package — a deductive estimate (×2 dies vs ~1,800 W for Rubin R100), not confirmed by a Nvidia primary source. [S7]
NVL576 “Kyber” System: 600 kW per Rack
The Kyber rack is a radical form factor change.
Configuration [roadmap vendor claim]: 576 GPU dies (144 packages × 4 dies) + 144 Vera CPUs.
Performance [roadmap vendor claim]:
- 15 EFLOPS FP4 inference (×21 vs current GB200 NVL72)
- 5 EFLOPS FP8 training
- 365 TB of “fast memory” system-wide (147 TB HBM4E + 218 TB LPDDR), 4.6 PB/s aggregate bandwidth
Physical architecture: 90° rotation of compute trays into blade form factor. 4 canisters × 18 compute blades. A PCB backplane replaces copper NVLink cables (lower latency, higher density). Power consumption: 600+ kW per rack [roadmap vendor claim].
In plain terms: 600 kW per rack is the power consumption of roughly 600 European households concentrated into a few square meters. For comparison, a standard datacenter rack draws 5 to 20 kW. Kyber is not a rack — it is an infrastructure constraint requiring purpose-built facilities.
NVL1152 variant: SemiAnalysis mentions a NVL1152 version (288 packages, 1,152 dies) “in development” [S5]. Not officially confirmed.
Sources: GTC 2025 keynote Jensen Huang (March 2025) [S3]; NextPlatform (March 2025) [S7]; SemiAnalysis “NVIDIA GTC 2025 — Built For Reasoning” (March 2025) [S5]; VideoCardz “NVIDIA unveils Rubin Ultra with 1TB HBM4e memory for 2027, Feynman architecture in 2028”; TrendForce (March 2025).
Feynman — 2028
Feynman (named after physicist Richard Feynman) is at this stage essentially a codename with architectural directions, not a specification. The associated CPU is named Rosa (named after Rosalyn Sussman Yalow, Nobel Prize in Physiology 1977 — Nvidia continues its tradition of naming CPUs after female scientists).
Available information comes from two primary sources: GTC 2025 (codename and initial directions) and GTC 2026 (3D stacking details, Rosa, NVLink 8 CPO).
Major Innovation: Vertical 3D Die Stacking [Partially Announced]
All previous generations assemble their dies in 2.5D (side by side on an interposer). Feynman introduces vertical 3D stacking of GPU dies — a first for processors at this power class. Vertical stacking reduces inter-die communication distances and enables further densification.
The thermal challenge is open: dissipating heat from vertically stacked dies with package power likely exceeding 1,000 W is an unsolved engineering problem as of GTC 2026. Nvidia acknowledged this challenge without communicating a specific solution.
Custom HBM Memory [Partially Announced]
For the first time, Nvidia is designing its own memory control logic integrated into the HBM stacks — cHBM (custom HBM). This breaks from the JEDEC standardized model: Nvidia no longer relies on standard specifications but co-designs memory with its suppliers. Likely HBM5 or enhanced HBM4E+. No capacity or bandwidth figures have been announced.
NVLink 8 CPO: Co-Packaged Optics [Partially Announced]
Feynman introduces NVLink 8 CPO (Co-Packaged Optics) — the first co-packaged optical integration in the NVLink standard. Scale-up and scale-out interconnects transition to optics directly integrated into the package, eliminating the losses of an external electro-optical connector. The rest of the network stack: BlueField-5 DPU, CX10 optical InfiniBand (3.2 Tb/s per port), Spectrum 7 Ethernet (204 Tb/s per switch), 8th-generation NVSwitch.
Rosa CPU [Partially Announced]
Rosa is described by Nvidia as a complete redesign of Vera, not an incremental evolution. No core/frequency/memory specifications are available.
What We Don’t Know
Feynman performance: no PFLOPS/EFLOPS figures officially announced. SemiAnalysis estimates 5 to 20× improvement vs Rubin [third-party estimate, speculative]. Process node: TSMC N2 or N2P probable (N2P volume production expected H2 2026, timeline compatible with 2028), not confirmed by Nvidia. Price: unknown.
In plain terms: Feynman is 2028 seen from 2026. We know the physicist’s name, the direction (3D stacking, optics, custom memory), and very few figures. Epistemic caution is warranted.
Sources: GTC 2025 keynote (March 2025) [S3]; GTC 2026 keynote Jensen Huang (March 2026) [S4]; Tweaktown “NVIDIA updates roadmap, with new details on its next-gen GPU Feynman coming in 2028” (March 2026); Tom’s Hardware “Nvidia updates data center roadmap with Rosa CPU and stacked Feynman GPUs” (March 2026) [S9]; WCCFTech “NVIDIA Feynman GPU Gets 3D Die-Stacking, Custom HBM, & Next-Gen Rosa CPU” (2026).
Implications for Datacenters and the Industry
TDP Progression: Liquid Cooling as the Only Option
Per-rack power density follows an exponential curve:
| Generation | GPU TDP | Reference Rack | Rack Power |
|---|---|---|---|
| H100 SXM (Hopper) | 700 W | DGX H100 8-GPU | ~10 kW |
| B200 (Blackwell) | 1,000 W | GB200 NVL72 | ~120 kW |
| B300 (Blackwell Ultra) | 1,400 W | GB300 NVL72 | ~170 kW [estimated] |
| R100 (Vera Rubin) | ~1,800 W [vendor claim] | NVL144 | ~300 kW [estimated] |
| R300 (Rubin Ultra) | ~3,600 W/package [estimated] | NVL576 Kyber | 600+ kW [vendor claim] |
Direct liquid cooling has been mandatory since Blackwell Ultra. The Kyber rack at 600+ kW has no air-cooled variant. Most existing datacenters are not equipped for densities above 30 kW per rack — Kyber requires specific infrastructure: liquid loops, hydraulic plumbing, adapted electrical distribution.
In plain terms: deploying a Kyber rack in 2027 is like trying to power a city block from a residential outlet. The infrastructure must be redesigned before the chips are even ordered.
TSMC CoWoS Packaging: A Capacity Bottleneck
CoWoS (Chip-on-Wafer-on-Substrate) packaging from TSMC is the interposer technology that allows assembling multiple dies on a single substrate. CoWoS capacity has been the real production bottleneck for Nvidia since 2024 — not the silicon itself.
Nvidia has secured ~60% of TSMC’s advanced packaging capacity for 2026. TSMC committed $56 billion in capex to double its CoWoS capacity in anticipation of the Rubin ramp. This lock-up means other customers (AMD, Broadcom, Google for its TPUs) must share the remaining 40%.
The CUDA Continuity Argument
Nvidia maintains backward CUDA compatibility across all generations: CUDA kernels written for Hopper run on Blackwell and will run on Rubin. This guarantee is the central strategic argument against migrating to AMD ROCm or hyperscaler ASICs — which cannot offer this software continuity across radically different architectures.
Dynamo: The Software Ecosystem That Precedes the Hardware
Announced at GTC 2025, Dynamo is a multi-GPU orchestration layer for reasoning models (smart token router for prefill/decode, auto-scaling GPU planner, NCCL optimized with 4× lower latency on small messages, KV-cache offload to NVMe). Dynamo is hardware-independent — it prepares the software ecosystem to exploit NVL144 and NVL576 before these systems are even available.
Pricing and Allocation: Structural Unknowns
No official pricing has been published for B300, Rubin, or Rubin Ultra. Analyst estimates:
- B300: $40,000–50,000 USD/GPU (H100 ~$37,000, B200 ~$40,000)
- Rubin: likely >$60,000 USD/GPU (HBM4 cost + TSMC N3P)
Allocation remains a constraint: TSMC CoWoS-L is the packaging bottleneck. Hyperscalers (Microsoft, Google, Meta, Amazon) hold pre-order contracts. Enterprises outside the top tier are expected to face 12 to 18 months of delays post-general availability.
Which Generation for Which Decision?
| Purchase Context | Relevant Generation | Why |
|---|---|---|
| Production deployment now (2025) | B300 Blackwell Ultra | Available H2 2025, +50% memory vs B200, for training and long-context inference. |
| 2026 cluster planning | R100 Vera Rubin | Full architectural leap, ×5 inference vs Blackwell, HBM4 — but HBM4 supply concentrated at SK Hynix. |
| 2027+ datacenter roadmap | R300 Rubin Ultra | Double the GPU silicon, 600 kW/rack — infrastructure must be designed now for a 2027 deployment. |
| Long-term outlook | Feynman 2028 | Confirmed directions (3D stacking, optics), specs unknown — no procurement decision possible. |
| Tight budget | Current Blackwell B200 | Available stock, MLPerf-documented performance, no premium paid for undelivered roadmap. |
Heuristic rules:
- Don’t pay for roadmap before delivery: all Rubin Ultra and Feynman specs are [roadmap vendor claim]. Blackwell (current) was delayed ~6 months in 2024 due to rack thermal issues. Delays are possible.
- Verify the facility before the chips: 300+ kW for NVL144, 600+ kW for Kyber — the constraint is the room, not the GPU budget.
- Anticipate HBM4 supply: only two suppliers (SK Hynix 70%, Samsung 30%). A production incident can shift the entire roadmap.
- CUDA continuity remains the strongest argument: for teams that have invested in GPU code, Nvidia’s backward compatibility is a documented and unchallenged claim versus alternatives.
Key Takeaways for 2026
Nvidia’s 2025-2028 roadmap is a structural response to accelerating in-house chip development by hyperscalers. Its coherence rests on three bets: (1) TSMC delivers the promised CoWoS capacity, (2) SK Hynix and Samsung deliver HBM4 and HBM4E on schedule, (3) customer datacenters invest in infrastructure capable of absorbing 600+ kW per rack. Across the four generations, two are delivered or in production (B300, R100), one is announced with partial specs (R300), and one is a codename with directional characteristics (Feynman). The prospective angle of this article will be updated as real deliveries and independent MLPerf benchmarks validate — or invalidate — vendor roadmap claims.