In brief

Modern GPUs have thousands of compute cores capable of executing billions of operations per second. Yet when a large language model generates text, those cores spend most of their time waiting — not computing. The bottleneck lies in memory: how much data can be read from it per second. High Bandwidth Memory (HBM) is the industry’s answer to this problem. Understanding how it works means understanding why AI datacenters cost so much and why chip supply remains so tight.


The memory wall: a 1994 problem that has only gotten worse

In 1995, researchers Wulf and McKee published a paper with a telling title: Hitting the Memory Wall: Implications of the Obvious. Their observation: processor compute speed grows far faster than memory access speed. The two curves diverge. They called this imbalance the memory wall.

Thirty years later, the problem has worsened dramatically. A high-end GPU can perform several hundred teraflops. But if data doesn’t arrive fast enough from memory, that compute power goes unused.

For large language models, the problem is particularly acute. Generating a token — a syllable, a word — requires loading the entire model weights from memory. For a 70-billion-parameter model stored in half-precision (FP16), that’s roughly 140 gigabytes per step. At an H100’s read speed (3.35 TB/s), that single operation takes ~42 milliseconds — regardless of available compute power. It is a hard limit, independent of core count.

Researchers characterize this phenomenon using the Roofline model: for large models in low-parallelism inference, the ratio of operations to bytes read is systematically below the threshold at which compute becomes the limiting factor. Memory bandwidth, not compute cores, bottlenecks inference.

In short: generating one word with a 70-billion model = re-reading 140 GB of weights, regardless of compute power. At 3.35 TB/s (H100), those 140 GB take 42 ms minimum — that is a physical floor, not an optimization defect.


What HBM is, and why it changes the rules

Conventional GPU memory (GDDR6X) uses a 384-bit bus. HBM takes a radically different approach: it stacks multiple layers of DRAM vertically in 3D, connects them via connections passing through the silicon (through-silicon vias, or TSVs), then places that stack directly alongside the processor on a silicon interposer.

The result: a 1024-bit bus per stack instead of 384. Bandwidth increases without needing to raise clock frequency — which would limit power consumption and heat dissipation.

The progression across generations is notable:

GenerationBandwidth per stackCapacity per stackExample GPU
HBM2e (2019)460 GB/sup to 16 GBA100 (Nvidia), MI200 (AMD)
HBM3 (2022)819 GB/sup to 24 GBH100 SXM5 (Nvidia)
HBM3e (2024)1.2 TB/sup to 36 GBH200 (Nvidia), MI300X (AMD), B200 (Nvidia)
HBM4 (planned 2026)~1.6 TB/sin development—

The H200 carries 141 GB of HBM3e for 4.8 TB/s total bandwidth — 43% more than the H100. AMD’s MI300X pushes to 192 GB and 5.3 TB/s. Nvidia’s Blackwell B200 exceeds 8 TB/s.

HBM also has one limitation: power consumption. An H100 SXM5 consumes 700 watts total, largely due to memory transactions. In AI datacenters, this thermal constraint governs installation density and operating cost — costs rarely highlighted in manufacturer spec sheets.

In short: HBM does not increase memory frequency, it widens its pipe. Instead of 384-bit buses, you have 1024 bits per stack, multiple stacks per GPU, and everything is glued directly next to the processor on a silicon interposer. This multiplies bandwidth by 5 without overheating the chip.


The shortage: it’s the memory, not the processor

One direct consequence of this architecture: HBM is difficult to manufacture. 3D stacking and TSVs require far more demanding fabrication processes than standard lithography. Yields are low, capacity ramp is slow.

Three players control virtually the entire global market:

  • SK Hynix (~50% market share): exclusive supplier to Nvidia for the H100 at launch, then primary supplier for subsequent generations.
  • Samsung (~40%): supplier to AMD and Google (TPUs), but encountered difficulties getting its HBM3e qualified by Nvidia in 2024.
  • Micron (~10%): entered HBM3e production in 2024, qualified for the H200.

This concentration creates structural dependency. When AI datacenter demand exploded from 2023 onward, production capacity did not keep pace. Nvidia had to ration its H100s, delivery lead times reached 12 to 18 months, and secondary market prices peaked around $40,000 — versus ~$30,000 list price. The premium directly reflected the scarcity of the memory component, not the processor.

What this shortage reveals: what costs the most in a high-end AI accelerator is the memory bonded to it.

The constraint does not lie in lithography — HBM DRAM is manufactured at advanced nodes (1Z nm). It lies in the TSV process and 3D assembly, which are precise operations, difficult to parallelize and to scale. SK Hynix announced $15 billion in investment over five years to expand capacity. The lag between investment and available production capacity exceeds two years. This is why analysts disagree on how quickly the tension will resolve: new capacity is coming, but demand driven by ever-larger models may absorb gains as fast as they arrive.


What this means for deployment

Understanding the memory wall has direct consequences for how models are deployed.

Batch size matters as much as bandwidth. When multiple requests are processed simultaneously (batching), model weights are loaded once to serve the entire batch — diluting the memory cost per request. A 70B model served in a batch of 64 requests costs roughly 64 times less in bandwidth per request than in single-request mode. This is why cloud inference services maximize request batching: it is a memory optimization before being a compute optimization.

Model size directly determines the number of GPUs required. A 70B model in FP16 occupies ~140 GB. An H100 offers 80 GB of HBM. At minimum, two H100s are needed to host this model, plus headroom for the KV cache. For models of 400B or more, this means configurations of 8 or 16 GPUs with tensor parallelism — each GPU holds only a fraction of the weights, and activations travel between them via NVLink. Inter-GPU bandwidth then becomes a second bottleneck.

Quantization is a hardware decision as much as a quality decision. Moving from FP16 to INT8 halves both the memory footprint and the volume to read per token. For single-GPU deployment, this is sometimes the difference between possible and impossible.

In short: three main levers to absorb the memory constraint — batching (serve 64 requests for the price of one weight read), quantizing (FP16 → INT8 = half the bytes), splitting (tensor parallelism over 8 GPUs). Cloud inference services use all three at once.


Memory as the only lever? Some important nuances

The thesis that “memory bandwidth is the limiting factor” is the most widespread, but it is not without counterarguments.

Software optimization can reduce the amount of data to transfer. Quantization techniques — reducing weight precision from 16 bits to 8 or 4 bits — mechanically divide the memory volume to read. Work by Dettmers et al. (LLM.int8(), NeurIPS 2022) and 4-bit quantization methods (GPTQ, AWQ) show that it is sometimes more efficient to compress data than to increase bandwidth. The race for terabytes per second is not the only answer.

Alternative architectures bypass the problem differently. Cerebras, with its wafer-scale chips, integrates memory directly on the chip as SRAM, achieving intrinsic bandwidths of several tens of TB/s — but with more limited capacity and high cost. This approach eliminates the memory wall by removing the physical distance between compute and memory.

Processing-in-memory (PIM) is a third path being explored: integrating compute units inside the memory stack itself. Samsung presented an HBM-PIM at IEEE Hot Chips 2022 with measured gains on certain ML operations. Generalization to production remains limited.

KV cache management is a challenge specific to long contexts. Work by Kwon et al. (PagedAttention, SOSP 2023) shows that a 128,000-token context for a 70B model can require tens of gigabytes of attention cache alone — a memory cost added on top of weights, forcing tradeoffs between batch size and accepted context length. The FlexGen paper (Sheng et al., ICML 2023) explores another approach: offloading part of memory to NVMe storage or CPU RAM, at the cost of latency, to enable inference on less expensive hardware.

Non-volatile memories via CXL constitute a fourth path, still experimental. The CXL (Compute Express Link) protocol would allow adding DRAM or persistent memory directly on the PCIe bus at low latency, creating an intermediate tier between fast HBM and slow storage. In 2025–2026, the viability of this tier in production for LLM workloads remains to be demonstrated.


Key takeaways

  • Large language models in inference are limited by memory bandwidth, not raw compute power: loading weights from memory at each generated token is the dominant constraint.
  • HBM partially solves this problem by stacking DRAM in 3D, connected via through-silicon vias, to reach 1024-bit buses — versus 384 bits for conventional memories.
  • Three companies (SK Hynix, Samsung, Micron) control global production, with processes difficult to scale quickly. The 2023–2024 shortage was a memory constraint, not a compute constraint.
  • Alternatives exist — quantization, wafer-scale architectures, processing-in-memory — that reduce the importance of HBM bandwidth without eliminating it.
  • HBM4 (planned 2026, interface widened to 2048 bits) targets 1.6 TB/s per stack, but the long-term trend remains unfavorable: compute power grows faster than memory capacity, and production constraints (TSV, 3D assembly) slow volume ramp.
  • Alternative paths — non-volatile memories via CXL, NVMe offloading — remain experimental in production for high-performance LLM workloads.