In brief

TPUs (Tensor Processing Units) are chips designed by Google specifically for neural network computations. Introduced in 2015, deployed in production in 2016, they are now in their seventh generation. Their architecture — the systolic array — differs fundamentally from GPUs. This difference explains both their exceptional performance under Google’s conditions and the significant friction for external users.

In short: a TPU is not a “Google GPU” — it is a chip from a different family, designed to do a single operation (matrix multiplication) as fast as possible. Where the Nvidia GPU is a Swiss Army knife capable of everything (graphics, physics simulation, AI), the TPU is an industrial steam iron: excellent at its job, unsuited for anything else. This extreme specialization is at once its strength (speed, cost, energy) and its limit (rigidity, JAX dependency, friction for anyone arriving with PyTorch code).


Why Google built its own chips

In 2013, Google engineers found that widespread adoption of neural networks for voice recognition would double their datacenters’ compute capacity requirements. Purchasing enough Nvidia GPUs to absorb this load represented a prohibitive cost. The decision was made: design a dedicated chip, optimized for a single operation — matrix multiplication.

The result is the TPU v1, deployed in 2016. Its goal is not to be versatile: it is built for inference, and nothing else. This radical specialization is the throughline of the entire TPU family.

The analogy that explains the architecture

A GPU resembles a multi-task workshop: hundreds of independent cores, each capable of executing different instructions, coordinated by a scheduler. Flexible, general-purpose, suited to a wide variety of workloads.

A TPU functions like an industrial assembly line. At the heart of the device: the systolic array, a grid of 256×256 multiply-accumulate units (MACs). Data flows through it in a synchronized stream, from cell to cell, without ever needing to be reloaded from memory. Each cell receives a partial result, enriches it, and passes it to the next.

This organization eliminates nearly all the memory accesses that slow down GPUs on dense matrix operations. When the workload is predictable and repetitive — as is a forward pass in a transformer — the assembly line outperforms the multi-task workshop by a wide margin.

The trade-off is symmetrical: if the workload deviates from the expected pattern, the assembly line seizes up. Irregular operations, dynamic compute graphs, unoptimized kernels — anything that breaks the flow — result in severe underutilization of the hardware.

In short: for the numerical version. An Nvidia A100 GPU typically reaches 60-80% hardware utilization on a standard ML workload. A TPU v4 on the same poorly calibrated workload can drop to 20-30%. Conversely, a TPU v4 well-optimized for its workloads (Gemini, PaLM) runs above 95% — a score nearly unreachable on GPU. The TPU is therefore a bet: if you accept rewriting code in JAX and embracing the systolic-array constraints, you recover a 2-5× gain on certain workloads. Otherwise, you pay for a GPU without getting its benefits.

A timeline in seven generations

TPU v1 (2016) — Inference only. int8 precision. 92 TOPS at 40 W. Jouppi et al. (2017, ISCA) measure gains of 15 to 30× on inference benchmarks versus contemporary CPUs and GPUs for Google production workloads.

TPU v2 (2017) — First generation capable of training. bfloat16 precision, the numerical format that Google would subsequently popularize across the industry. 45 TFLOPS. Introduction of pods: 64 interconnected chips.

TPU v3 (2018) — 420 TFLOPS. Liquid cooling introduced to handle thermal density. Pods of 1,024 chips.

TPU v4 (2021-2022) — Two structural innovations. The OCS (Optical Circuit Switch) enables dynamic reconfiguration of the network topology between chips via optical connections — an industry first. SparseCores accelerate operations on embeddings (sparse lookup tables), with gains of 5 to 7× on these specific workloads. Overall performance: 2.1× vs v3, 1.2 to 1.7× vs A100 according to Jouppi et al. (2023, ISCA). PaLM training mobilizes 6,144 TPU v4 chips.

TPU v5e / v5p (2023) — Two distinct profiles. The v5e targets inference efficiency. The v5p targets training scale: it is the chip used for Gemini Ultra.

Trillium v6 (2024) — 4.7× the performance of the v5e. The systolic array remains at 256×256 MACs. The improvement comes primarily from clock speed and memory.

Ironwood v7 (2025) — A bet on pure inference. 4,614 TOPS. 192 GB of HBM3 per chip. Pods of 9,216 chips. The announced absence of continuous training capabilities on this generation reflects a strategic choice: Google anticipates that production workloads will remain dominated by large-scale inference.

In short: the seven-generation trajectory in raw numbers. v1 (2016): 92 TOPS. v7 Ironwood (2025): 4,614 TOPS — that is ×50 in performance per chip in nine years. Over the same interval, pods went from 64-chip configurations (v2) to 9,216 chips (v7), that is ×144 in size. The Exaflop threshold — one quintillion operations per second — was crossed around v4; v7 reaches 42.5 Exaflops per superpod, comparable to the ~13 Exaflops of the Frontier supercomputer (USA, 2022). Google now has internal compute capacity comparable to the largest national supercomputers, exclusively for AI.

Scale-out: topology as advantage

An isolated TPU is already powerful. But Google’s systemic advantage lies in networking. TPU v4 chips and later communicate via the ICI (Inter-Chip Interconnect), organized as a 3D torus. This topology enables very high bandwidth between nearby chips, and controlled latency across configurations of several thousand chips.

Multislice, introduced for recent generations, allows a pod to be split into independent slices for jobs of different sizes, and dynamically recomposed. PaLM and Gemini could not have been trained without this infrastructure.

Nvidia offers comparable solutions — NVLink, NVSwitch, InfiniBand — but Google’s vertical integration (chips + network + datacenter + operating system) provides an end-to-end advantage that is difficult to replicate outside this context.

The tension: internal power, external friction

This is where the picture becomes complicated. The performance figures published by Google are real — but they are achieved under conditions that the vast majority of users will not reproduce.

The JAX/XLA ecosystem

TPUs are designed for JAX and its XLA compiler. JAX is Google’s research framework: performant, elegant for linear algebra, but with a steep learning curve for anyone coming from PyTorch. XLA compiles compute graphs and optimizes them for the target hardware — it is what exploits the systolic array.

PyTorch is supported via torch_xla, but this abstraction layer introduces constraints. Some operators are not natively supported. FlashAttention — the attention optimization that has become an industry standard — took several years to be ported effectively to TPU, whereas it was available on GPU from its earliest versions. The dependency on JAX is not a configuration detail: it is a structural prerequisite.

The access asymmetry

Google uses TPUs internally without going through Cloud TPU. The configurations deployed for PaLM or Gemini are not available for purchase. What Google sells on Cloud TPU is a version of the internal infrastructure, with limitations on topology, allocation duration, and configuration surface.

For a lab wanting to train a 70 billion parameter model on TPU v5p, the availability of pods, managing preemptions, and network configuration represent a significant operational investment — whereas H100 GPUs on AWS or Azure fit into better-documented and more portable workflows.

CUDA vs XLA as ecosystem divide

CUDA is not merely a framework: it is twenty years of accumulated optimizations, third-party libraries, tutorials, and profilers. XLA is younger, less well-tooled for debugging, and the community surrounding it is more restricted. The majority of research papers publish PyTorch/CUDA code. Reproducing these results on TPU requires non-trivial porting effort.

This divide tends to narrow — Google invests in documentation and tooling, and JAX is gaining traction in certain research communities. But in 2026, the gap remains real.

TPU vs Nvidia GPU: when to choose what

ContextRecommendationWhy
Frontier model training in-house (Google scale)TPU v5p or IronwoodCost ~2× lower vs Nvidia GPU, energy -67%, ICI 3D-torus infrastructure to scale to 9,000+ chips. Available exclusively at Google.
Standard academic research, reproducible papersNvidia GPU (A100/H100)Mature CUDA ecosystem, FlashAttention and libraries available immediately, portable PyTorch code with no effort.
Cloud production with multi-provider portability requirementsNvidia GPUAvailable on AWS, Azure, GCP — avoid Cloud TPU lock-in. Transferable workload.
Large-scale inference for Google application (Gemini, Search)Cloud TPU v5e or IronwoodIronwood v7 designed for pure inference — 192 GB HBM3 per chip, bfloat16 optimizations.
Starting an ML project, fast prototypeNvidia GPUDocumentation, tutorials, Stack Overflow community, profilers — everything is mature and accessible.
Already JAX pipeline in production, experienced XLA teamTPU justifiedThe learning overhead is already paid — exploiting the TPU perf/cost gain becomes coherent.

Key takeaways

  • TPUs are built on the systolic array: a synchronized data flow that eliminates memory round-trips on dense matrix operations. Highly effective for predictable workloads, rigid for others.
  • Seven generations in ten years, with measured gains from 15-30× (v1 vs CPU/GPU 2016) to 4,614 TOPS (Ironwood v7, 2025). Each generation has brought a structural innovation: bfloat16 (v2), pods (v2), OCS (v4), SparseCores (v4).
  • Google’s real advantage is not an isolated chip — it is vertical integration: chip + ICI 3D torus interconnect + datacenter + XLA compiler.
  • The ecosystem remains centered on JAX/XLA. PyTorch is supported but with friction. FlashAttention and other standard optimizations were long unavailable or underperforming on TPU.
  • The majority of published performance figures are achieved under internal conditions inaccessible via Cloud TPU. For external users, the performance-to-operational-effort ratio remains lower than Google’s benchmarks suggest.