At a glance

In 2026, two AI chips dominate the top of the accelerator market: Google’s TPU v7 Ironwood, announced in April 2025, and Nvidia’s GB200 NVL72, available in production since late 2024. This is the first time since the H100 (2022) that both leaders are launching their flagship generations in the same time window — making an unprecedented head-to-head comparison possible with contemporary figures.

The result of this comparison is not a single winner. The two systems target different problems: Ironwood bets on inference at very large scale in pods of 9,216 chips; GB200 NVL72 optimizes flexible training in units of 72 GPUs. The choice depends on use case, mastered software ecosystem, and hardware access.

This article focuses on the 2025-2026 head-to-head. For a detailed profile of each actor, see Google TPU and Nvidia GPU and CUDA.


Understanding the problem: why 2026 is the year of the comparison

Imagine two roads built to carry different types of cargo. The first is a multi-lane highway, designed for regular, predictable convoys — it’s fast but cannot turn around. The second is a versatile road network, less optimized for any single vehicle type, but capable of adapting to any route. TPUs are the first road, GPUs the second.

Two roadmaps converge simultaneously in 2025-2026:

  • On the Nvidia side: the Blackwell architecture (B100, B200, GB200 NVL72) has been available in production since late 2024. The GB200 NVL72 — 72 B200 GPUs interconnected in a 120 kW liquid-cooled rack — is the reference system for frontier model training.
  • On the Google side: TPU v7 Ironwood announced in April 2025, general availability expected late 2025 / early 2026. The first TPU explicitly designed for “the inference era” (source: blog.google, April 2025).

An external catalyst gives this head-to-head a concrete dimension: Gemini 2.5 is trained entirely on TPU (v5p according to SemiAnalysis); Anthropic signed a massive contract for ~400,000 TPU v7 purchased and 600,000 leased via GCP, targeting ~1 GW of compute capacity (source: SemiAnalysis 2025). These industrial-scale deployment decisions, made by actors who compared both options, are strong signals about the respective value of each architecture.

In plain terms: for the first time since 2022, Google and Nvidia have comparable flagship chips at the same date. This synchrony enables a head-to-head on contemporary figures rather than on staggered generations.


Architectures — Systolic Array versus SIMT

TPU Ironwood: the synchronized dual-chiplet flow

The heart of the TPU is the systolic array: a grid of MACs (multiply-accumulate units) where data flows from cell to cell without memory reloading between operations. Think of an assembly line where each worker receives a part from their neighbor, does their work, and passes it on — without ever fetching materials themselves.

TPU v7 Ironwood takes an architectural step by adopting a dual-chiplet design: 2 chiplets per chip, each integrating a TensorCore, two SparseCores and 96 GB of HBM3e. The die-to-die (D2D) bus linking them is 6× faster than a standard ICI (inter-chip interconnect) link (source: docs.cloud.google.com).

SparseCores (introduced in TPU v4) handle operations on sparse embeddings — lookup tables for recommendation models — with measured gains of 5-7× on these specific workloads (Jouppi et al., ISCA 2023).

Vocabulary

  • Systolic array: a grid of simple processors where data flows from neighbor to neighbor without central memory access. Very efficient on predictable flows.
  • MACs (multiply-accumulate units): the basic operation in neural networks — multiply two numbers and add the result to an accumulator.
  • HBM3e (High Bandwidth Memory): type of memory stacked vertically on the chip, offering 5-10× higher bandwidth than conventional DDR.
  • ICI (Inter-Chip Interconnect): high-speed link connecting TPU chips together within a pod.

In plain terms: the TPU Ironwood is an ultra-specialized factory — it processes regular compute flows at record speed, but is less comfortable with dynamic computation graphs or irregular operations.

GPU GB200: the flexible network of 72 processors

The GPU (Single Instruction Multiple Threads) executes thousands of threads in parallel, each capable of taking different execution paths. This flexibility is the foundation of the CUDA ecosystem — unlike the TPU highway, the GPU road network can adapt its route to each request.

The Blackwell B200 adopts a multi-die architecture: two chips linked by a 10 TB/s chip-to-chip internal bus. The GB200 NVL72 connects 72 B200s via NVLink 5.0 (1.8 TB/s) in an all-to-all topology in a 120 kW liquid-cooled rack (source: nvidia.com/gb200-nvl72, 2024).

The Transformer Engine (integrated since Hopper, version 2 on Blackwell) is a dedicated unit that automatically switches between FP8, FP4 and higher precisions according to the needs of each model layer. This is a dynamic adaptation mechanism that the TPU’s systolic array, fixed by nature, cannot reproduce in the same way.

In plain terms: the GB200 NVL72 is a network of 72 versatile processors capable of adapting to varied model architectures. Its flexibility has a cost: it is not as optimized as the TPU on predictable flows, but it excels where workloads change frequently.

This is where the two systems diverge most strongly, and where the architecture choice becomes strategic.

DimensionTPU v7 Ironwood (pod)Nvidia GB200 NVL72
Intra-domain topology3D torus (ICI)All-to-all NVLink 5.0
Bandwidth per link1.2 TB/s bidir (ICI)1.8 TB/s (NVLink 5.0)
Max intra-domain size9,216 chips (1 pod)72 GPUs (NVL72)
Inter-node extensionJupiter network (reconfigurable optical OCS)InfiniBand NDR/XDR
Topology reconfigurabilityOCS: dynamically reconfigurableFixed (NVSwitch fabric)

NVLink offers 1.8 TB/s per link versus 1.2 TB/s ICI — Nvidia has the edge per individual link. But Google connects 9,216 chips in a single coherent domain, versus 72 GPUs for the NVL72. For very large models (beyond 1 trillion parameters), the TPU topology enables combinations of parallelism — data, sequence, tensor, pipeline — that are physically inaccessible on an NVL72 (source: SemiAnalysis, “TPUv7 The 900lb Gorilla”, 2025).

The aggregate bandwidth of a full Ironwood pod (9,216 chips) reaches 42.5 ExaFLOPS in FP8 (source: blog.google, April 2025). An equivalent Nvidia GPU pod would involve ~128 NVL72 units, connected by InfiniBand — with the associated latencies and topology management challenges.

In plain terms: Nvidia wins on per-link bandwidth, Google wins on maximum scale within a coherent domain. For training a 70B parameter model, both systems work well. For training a 10-trillion parameter model, the TPU pod offers interconnect coherence that NVL72 cannot reproduce without additional InfiniBand infrastructure.


Roadmaps 2024-2026 — Generation by generation

Google TPU: from v5 to Ironwood

GenerationYearFP8 TOPS/chipHBM (GB)Mem BW (TB/s)Note
TPU v5e2023~19716-321.6Standard inference efficiency
TPU v5p2023~459954.6Training scale (Gemini Ultra)
Trillium (v6e)2024~9181444.2-4.84.7× vs v5e perf/chip
Ironwood (v7)20254,6141927.4Dual-chiplet; inference-first

Sources: docs.cloud.google.com, cloud.google.com/blog (Trillium GA), blog.google (Ironwood April 2025).

The Ironwood leap is spectacular: the progression from TPU v5e to v7 represents ~23× in FP8 TOPS, 6× in HBM and 4.6× in memory bandwidth. Google claims 2× perf/watt vs Trillium and 10× vs v5p [vendor claim — not independently confirmed at this stage].

September 2026 update — the eighth generation is announced, and it splits in two. Google drops the single chip and announces two distinct architectures: TPU 8t for training and TPU 8i for inference.

TPU 8t (training)TPU 8i (inference)
Scale9,600-chip superpod, 2 PB of shared HBM, 121 ExaFLOPS—
Memory—288 GB HBM, 384 MB on-chip SRAM (×3 vs Ironwood)
Interconnectinterchip bandwidth doubled vs Ironwood19.2 Tb/s, doubled, for MoE models
Claimed gain vs Ironwood~3× compute per pod+80% performance per dollar

Both chips are announced at “up to 2× better performance-per-watt” than Ironwood, with general availability “later this year”. [vendor claim — none of these figures is validated by a third-party benchmark to date]

In short: the notable fact is not the performance gain, it is the separation. After seven generations of a single chip, Google concedes that training and serving are no longer the same job — precisely the thesis Ironwood already carried by declaring itself “inference-first”, but now etched into two different pieces of silicon. Nvidia, meanwhile, keeps selling one chip that does both.

Nvidia: from H100 to GB200

GenerationYearFP8 TOPS/GPUHBM (GB)Mem BW (TB/s)Note
H100 (Hopper)20223,958803.35Transformer Engine FP8
H20020243,9581414.8+HBM3e vs H100
B200 (Blackwell)20244,5001928.0Multi-die, TE gen2
GB200 (Grace Blackwell)2024~5,0001928.0Integrated CPU+GPU

Sources: nvidia.com Blackwell whitepaper 2024, developer.nvidia.com.

2025 crossover point: Ironwood (4,614 FP8 TOPS) exceeds the standalone B200 (4,500 TOPS) but remains below the GB200 (~5,000 TOPS). These figures are all vendor-sourced and have not yet been validated by independent third-party benchmarks for Ironwood.

In plain terms: on chip-to-chip specs, Ironwood and B200 are in the same order of magnitude in 2025-2026. The difference plays out at scale (9,216 vs 72), in the software ecosystem, and in hardware access.


MLPerf Benchmarks — What independent data shows

MLPerf benchmarks (published by MLCommons) are the only independent comparative references in the sector. Their limitations are important to note: Google and Nvidia choose which configurations they submit, and results reflect software optimization as much as hardware capability.

MLPerf Training v4.1 — November 2024

First official round including Google Trillium (TPU v6e).

  • Google configuration submitted: 2× 256-pod Trillium (2,048 chips)
  • GPT-3 175B result: Trillium ~27.6 min vs TPU v5p ~29.6 min on 2,048 chips — ~8% gain
  • Trillium scaling efficiency: 99% (vs 94% for v5p) — the chip scales near-linearly (source: cloud.google.com Trillium MLPerf 4.1)
  • Trillium perf/dollar: 1.8× vs v5p (~45% training cost reduction according to Google) [vendor claim]

On this same round, Nvidia H100/H200 held the fastest absolute times at equivalent accelerator counts — Nvidia’s raw speed advantage remains structural for GPT-3 (source: xpu.pub analysis, November 2024).

MLPerf Training v5.0 — June 2025

Major change: GPT-3 175B is replaced by Llama 3.1 405B as the flagship benchmark, reflecting the evolution of real workloads toward large dense models.

Key results (source: MLCommons June 2025):

  • 201 results from 20 organizations — absolute participation record
  • Hardware present: AMD MI300X/MI325X, Nvidia B200 and GB200, Google Trillium
  • Nvidia is the only submitter present on all benchmarks, including Llama 3.1 405B
  • Inter-generation improvements (v4.1 → v5.0): +2.28× on Stable Diffusion 8-GPU, +2.10× on Llama 2 70B LoRA, +3.68× on Stable Diffusion multi-node

Important gap: granular TPU vs GPU figures on Llama 3.1 405B in MLPerf v5.0 were not publicly available at the time of writing. MLCommons confirmed Trillium’s presence but detailed performance figures had not yet been published at the level required for a direct comparison.

Ironwood gap: TPU v7 Ironwood has not yet submitted to MLPerf as of April 2026. The submission is expected in 2026. Any comparison of Ironwood vs GB200 therefore rests on vendor specs, not independent benchmarks.

MLPerf Inference v5.0 — April 2025

  • Nvidia Blackwell (B200/GB200) present with significant inter-generation gains (source: developer.nvidia.com)
  • Google TPU not listed in MLPerf Inference v5.0

For inference, the available independent reference is the Artificial Analysis benchmark (Llama 3.3 70B, vLLM framework):

AcceleratorCost/million tokensSource
H100 SXM~$1.06Artificial Analysis, May 2025
MI300X (AMD)~$2.24Artificial Analysis, May 2025
TPU v6e Trillium~$5.13Artificial Analysis, May 2025

Critical caveat: this 5× disadvantage figure for Trillium must be interpreted carefully. Lucas Beyer (Google Research) notes that Google does not serve its production workloads via vLLM-TPU — the vLLM framework has received far less optimization for TPU than for GPU. This benchmark measures “non-optimized vLLM on Trillium”, not the ceiling of the TPU architecture for inference. In production, Google uses its own inference stacks (JAX/MaxText on AI Hypercomputer), which are not measured in this benchmark.

In plain terms: available MLPerf benchmarks through mid-2026 show Trillium competing with Nvidia GPUs in large-scale training, with a slight scaling efficiency advantage. For inference, independent data remains limited — and vLLM comparisons structurally disadvantage TPUs whose inference ecosystem is not yet mature outside of Google.


Software Ecosystem — The real differentiator

If hardware specs are comparable, it is the software ecosystem that determines the choice for virtually all teams outside of Google.

CUDA: twenty years of accumulation that cannot be reproduced

CUDA represents a moat that goes well beyond hardware: 4 million trained developers, more than 3,000 compatible applications, highly optimized libraries (cuBLAS, cuDNN, NCCL, TensorRT). The major frameworks (PyTorch, TensorFlow, JAX) can all run on CUDA.

For an ML team in 2026, starting on GPU means: standard PyTorch code, available tutorials, 95% of published papers with reproducible CUDA/PyTorch code.

JAX/XLA: the native TPU path

JAX + XLA is the native execution path on TPU. XLA compilation fuses operators and optimizes memory accesses statically — unlike GPU CUDA where the JIT compiler optimizes dynamically.

Ironwood (TPU v7) breaks with TensorFlow: TensorFlow is not supported on v7. Only JAX and PyTorch/XLA are supported (source: docs.cloud.google.com/tpu/docs/tpu7x). For teams accustomed to TensorFlow, this represents an additional migration step.

Pallas (2024-2025) is Google’s new kernel language for TPU, analogous to Triton for GPU. It allows writing custom kernels directly in Python, reducing dependence on pre-compiled XLA kernels.

Ongoing convergence (2025-2026)

Two developments are reducing the software gap:

  1. vLLM-TPU redesign (October 2025): JAX becomes the unified lowering path for all vLLM models on TPU, whether defined in PyTorch or JAX (source: blog.vllm.ai, October 2025). This redesign significantly reduces the PyTorch→TPU friction for inference.

  2. PyTorch/XLA: PyTorch support on TPU improves session after session. Most standard PyTorch operations now work via this bridge.

Structural gap that persists: XLA:TPU compiler and MegaScaler (multi-pod orchestration) remain proprietary (SemiAnalysis 2025). Open-sourcing these components would accelerate external adoption — this choice remains in Google’s hands.

Software ecosystem summary

DimensionGoogle TPU (v7 Ironwood)Nvidia GPU (GB200)
Native frameworkJAX + XLACUDA + PyTorch
Secondary frameworkPyTorch/XLA (real friction)JAX, TensorFlow (also native)
Custom kernelsPallas (2024-2025)Triton / CUDA kernels
Code portabilityXLA hardware-agnostic (GPU+TPU)CUDA proprietary; Triton cross-GPU
Debug toolingImproving, less matureCUDA Profiler, NSight, very mature
CommunityMainly research/big tech4M+ developers, all sectors
Reproducible papersMinority publish in JAXAlmost all in PyTorch/CUDA

In plain terms: if your team works in PyTorch and follows academic papers, the GPU is the natural choice in 2026. If you work in the Google ecosystem (GCP, JAX, MaxText) or are ready to invest in the transition, Ironwood offers advantages at large scale.


Who uses what in 2026 — Real deployment decisions

Google: TPU everywhere, except for on-prem customers

  • Gemini 2.5: trained primarily on TPU v5p (source: SemiAnalysis 2025)
  • Gemini 3: Ironwood TPU according to IntuitionLabs (2025) [secondary source — not officially confirmed by Google]
  • Gemini inference: via AI Hypercomputer + JAX/MaxText, on TPU
  • Notable exception: Gemini 2.5 on Google Distributed Cloud (enterprise on-prem deployments) runs on Nvidia Blackwell GPU (source: Data Center Dynamics, wionews.com 2025). Economic logic: enterprise customers don’t have access to TPU pods — they buy deliverable hardware.

Frontier labs outside Google

  • Anthropic: massive contract signed for TPU v7 Ironwood (~1M chips total, purchased + leased via GCP) according to SemiAnalysis 2025. Partial pivot from AWS Trainium underway.
  • OpenAI: Nvidia exclusive (H100/B200). SemiAnalysis notes OpenAI obtained ~30% tariff reduction on its Nvidia fleet solely due to the TPU competitive threat — a sign that the negotiating leverage is real.
  • Meta: GPU (H100/H200/B200) + AMD MI300X for Llama 405B inference. No known TPU production deployment.
  • Mistral, Cohere, startups: GPU via cloud providers (H100/H200). No TPU.

Emerging usage patterns

Large-scale inference is the segment where custom ASICs (TPU, Trainium, Inferentia) have the clearest economic advantage: fixed topology, ability to specialize each component. Training of new architectures remains dominated by GPUs, thanks to CUDA flexibility and the research ecosystem (papers published in PyTorch/CUDA).


Cloud costs — What prices signal about positioning

AcceleratorOn-demand price/hourSource
H100 SXM~$2.69–3.00getdeploying.com (aggregator, April 2026)
H200~$3.50–4.50getdeploying.com
B200~$4.40–6.00getdeploying.com
GB200 (per GPU)~$10.50–21.53getdeploying.com
TPU v6e Trillium$2.70/chipcloud.google.com/tpu/pricing
TPU v7 IronwoodNot published (preview)[GAP — pricing not announced at this stage]

Google claims TPU training cost 40-60% lower than GPU [vendor claim — not independently confirmed] and 30-44% lower TCO than GB200 according to SemiAnalysis (paywalled source, partial access).

GB200 is the extreme case: up to $21.53/GPU/hour, or ~$1,550/hour for a full NVL72. This pricing reflects its positioning — frontier model training, not standard inference.

In plain terms: Trillium is the cheapest per chip in the public catalog. Ironwood has no announced pricing yet. GB200 is the most expensive per GPU but also the most capable for large training runs. The real cost/performance comparison depends on the workload and software optimization — vendor claims should be taken with caution.


Decision matrix — Choosing in 2026

ContextRecommendationWhy
Frontier model training >1T parameters, GCP accessTPU Ironwood pod9,216 chips in coherent domain offer interconnect topology inaccessible with NVL72; near-linear scaling confirmed in MLPerf.
Dense model training 7B–70B, PyTorch teamNvidia GB200 NVL72Native CUDA compatibility, 20 years of tooling, virtually all papers immediately reproducible.
Very large-scale inference, Google/JAX stackTPU Trillium or IronwoodMore competitive per-chip catalog cost; optimized for regular compute flows; requires JAX or PyTorch/XLA.
Standard inference, vLLM/PyTorch stackNvidia H100/H200vLLM better optimized on GPU; $1.06/M tokens H100 vs $5.13 Trillium (vLLM benchmark — caveat: TPU not optimized).
Research and reproducibility (papers, startups)Nvidia GPU95% of papers published in PyTorch/CUDA; ecosystem of tutorials and standard benchmarks.
On-premises enterpriseNvidia BlackwellTPU not available outside GCP; Nvidia deploys on Distributed Cloud and on-prem.
Mixed GPU+TPU team, multi-cloudEvaluate case by caseJAX is portable GPU+TPU; CUDA remains proprietary. PyTorch/XLA bridge reduces switching cost to TPU.

Practical heuristics

  1. If you are not already on GCP/JAX, the migration cost to TPU is high — the potential economic advantage must be calculated including software refactoring costs.
  2. Vendor perf/watt comparisons are not cross-comparable: Google claims 2× perf/watt for Ironwood vs Trillium [vendor claim]. Nvidia claims “25× H100 at same power” for GB200 [vendor claim]. These figures have not been validated by independent third parties.
  3. Ironwood’s absence from MLPerf (as of April 2026) is a real gap — any purchase decision on Ironwood rests on vendor specs without independent benchmark validation.
  4. For very large clusters, topology matters more than TOPS/chip: 9,216 chips in a coherent 3D torus versus 9,216 chips over InfiniBand are two fundamentally different parallelism environments.