In brief
When a server runs out of memory for a large language model, the first instinct is to buy more DRAM sticks. But the number of slots is limited, and beyond a certain threshold, the only option is to add an entire server. CXL (Compute Express Link) offers a third path: plugging additional memory modules into the server’s PCIe port — like plugging in a disk, but with performance close to local DRAM. Multiple servers can even share a common CXL memory pool through a switch. In 2025–2026, this technology is moving from the lab stage to early cloud instances, driven by one specific use case: storing the KV cache for long-context LLM inference.
The problem: the CPU server overflows before the GPU does
To understand why CXL exists, look at the CPU servers that orchestrate LLM inference — not just the GPU accelerators.
A modern x86 server (Intel Xeon, AMD EPYC) has 8 to 12 DDR5 memory channels, representing 8 to 24 physical DIMM slots depending on the platform. That is the mechanical ceiling. Filling all slots with 64 GB sticks reaches at most 1.5 TB of DRAM per socket. Beyond that, adding a complete server is the only option.
For LLM inference, this ceiling creates a structural imbalance: a 70B model’s weights in half-precision weigh 140 GB, and the KV cache (the attention keys and values computed for each context token) adds on top of those weights. For a 128,000-token context on a model of that size, the KV cache alone can exceed 50 GB per request. Multiplied across simultaneous requests, this quickly overflows even a well-dimensioned DRAM pool.
The traditional answer — add a server — duplicates the full cost: CPU, chassis, power supply, licenses. CXL proposes a decoupling: keep compute on the main server, move memory capacity to external modules connected via PCIe.
Vocabulary
- KV cache: memory structure storing intermediate attention results for each token already read. Its size grows linearly with context length.
- DRAM: standard server RAM (DDR4 or DDR5 sticks). Latency ~80 ns, bandwidth ~50 GB/s per channel.
- PCIe: standard interconnect bus between the CPU and peripherals (GPU, SSD, network cards). Universal in all x86 servers.
- Memory pooling: sharing a memory pool across multiple hosts, each host receiving a dedicated partition (or shared in CXL 3.0).
What CXL changes — and what it is not
Before diving into the mechanisms, a fundamental clarification: CXL is often confused with other memory technologies because they all address memory problems in AI systems. The three players are different.
HBM (High Bandwidth Memory) is 3D-stacked memory bonded directly onto the GPU die, connected by a 1024-bit bus. It is intra-chip, ultra-fast (up to 5.3 TB/s on the MI300X), but limited in capacity (192 GB maximum in 2024) and reserved for compute accelerators. HBM is not accessible from another server. Its dedicated article details this architecture.
NVLink and InfiniBand are compute interconnects between GPUs or between nodes. They synchronize training gradients, orchestrate tensor parallelism. Their role is to circulate compute data between processors, not to share memory in the sense of a coherent address space.
CXL operates differently: it is a memory coherence protocol built on PCIe. It allows the CPU to address external memory (a module plugged into PCIe) exactly as if it were local DRAM — with the same load/store semantics, without an intermediate network layer. The difference is fundamental: accessing CXL memory requires no NIC, no network stack, no software lock — just an ordinary read or write instruction.
In plain terms: HBM = ultra-fast memory bonded to the GPU. NVLink/InfiniBand = highways between GPUs. CXL = extending a server’s DRAM via the PCIe port, with possible sharing between servers.
CXL versions — a six-step ramp-up
CXL 1.0 / 1.1 (2019–2020) — simple expansion
First version, built on PCIe 5.0 (32 GT/s per lane). CXL defines three sub-protocols that activate depending on the device type:
- CXL.io: classic I/O transactions, inspired by PCIe.
- CXL.cache: the device can read and cache host memory (device→host coherence).
- CXL.mem: the host can read and write the device’s memory (memory expansion).
In version 1.1, only Type 1 devices (accelerators without their own memory) and Type 2 devices (accelerators with shareable memory) are supported. Intel Xeon Sapphire Rapids (January 2023) is the first mainstream CPU to implement CXL 1.1, but with a major constraint: no Type 3 support (pure memory modules). AMD EPYC Genoa (2022) integrates full Type 3 support, taking a one-year lead over Intel.
CXL 2.0 (2020) — pooling arrives
This is the version that makes memory pooling possible: multiple CPU hosts can connect to a CXL switch, itself connected to a pool of memory modules. Each host receives a dedicated partition of the pool (memory is not simultaneously shared, but can be dynamically redistributed).
Astera Labs launches Leo, the first commercial CXL switch ASIC, capable of resizing partitions without rebooting the host CPUs — critical for multi-tenant cloud environments. The introduction of Type 3 devices (pure memory modules, no compute) opens the door to the first commercial memory expanders.
CXL 3.0 (2022) and 3.1 (November 2023) — true sharing
Move to PCIe 6.0 (64 GT/s per lane, PAM-4 encoding). Bandwidth doubled. The fundamental innovation: memory sharing — multiple hosts simultaneously access the same bytes of memory with hardware coherence. In 3.0, the fabric supports multi-level topologies (rack-scale), not just a single switch.
CXL 3.1 (November 2023) extends pooling and sharing capabilities, introduces trusted compute environments for secure partition isolation, and optimizes resource management in multi-tenant environments.
CXL 3.2 (December 2024) and 4.0 (November 2025)
CXL 3.2 improves monitoring and memory device management, adds the Trusted Security Protocol (TSP) for cryptographic partition isolation. CXL 4.0, based on PCIe 7.0, doubles bandwidth again (128 GT/s) and improves reliability (RAS — Reliability, Availability, Serviceability).
| Version | Year | PCIe base | Bandwidth | Key feature |
|---|---|---|---|---|
| CXL 1.0/1.1 | 2019–2020 | PCIe 5.0 | 32 GT/s | Memory expansion Type 1/2 |
| CXL 2.0 | 2020 | PCIe 5.0 | 32 GT/s | Pooling + Type 3 + switching |
| CXL 3.0 | 2022 | PCIe 6.0 | 64 GT/s | True sharing + rack-scale fabric |
| CXL 3.1 | Nov. 2023 | PCIe 6.0 | 64 GT/s | Trusted environments + extended pooling |
| CXL 3.2 | Dec. 2024 | PCIe 6.0 | 64 GT/s | Monitoring + TSP security |
| CXL 4.0 | Nov. 2025 | PCIe 7.0 | 128 GT/s | Bundled ports + improved RAS |
Sources: CXL Consortium — BusinessWire, Nov. 2025 and Dec. 2024.
In plain terms: CXL 1.x = add memory to one server. CXL 2.0 = share a pool across multiple servers (dedicated partitions). CXL 3.0 = multiple servers read the same bytes simultaneously. Each generation builds on the previous without breaking changes.
The three CXL device types
CXL defines a taxonomy of devices based on which sub-protocols they activate. This taxonomy determines real-world use cases and adoption.
Type 1 — Accelerators without their own memory
These devices activate CXL.cache + CXL.io only. Typical examples: SmartNICs, intelligent network cards. The device can read and coherently cache host memory without exposing its own memory. Use case: network acceleration with direct access to host DRAM, zero-copy.
Type 2 — Accelerators with shareable memory
These devices activate all three sub-protocols: CXL.io + CXL.cache + CXL.mem. They have so-called HDM (Host-managed Device Memory): the CPU can read the device’s memory, and the device can read the host’s DRAM — within a shared coherence domain. GPUs, FPGAs, and compute ASICs theoretically fall into this category.
In practice, high-performance GPUs (H100, MI300X) do not use CXL — they favor their proprietary interconnects (NVLink) and integrated HBM for raw throughput. The debate on Type 2 GPU adoption in production remains open [NOT VERIFIED — SemiAnalysis position, partially paywalled source].
Type 3 — Memory expanders (the pure modules)
These devices activate only CXL.mem + CXL.io. No compute: they are DRAM enclosures connected via PCIe. The CPU addresses them exactly like regular sticks. This is the most commercially deployed type in 2024–2025, and the most relevant for LLM KV cache.
Available product examples:
- Samsung CMM-D: 512 GB DDR5, PCIe 5.0 x8 interface, proprietary CXL ASIC. Cumulative capacity up to 16 TB per CPU announced (Samsung marketing claim, actual performance to be validated). CMM-D 2.0 targeting CXL 3.1 sampling in late 2025.
- SK Hynix CMM-D: 96 GB DDR5, EDSFF ES.3 form factor, PCIe Gen 5 x8, CXL 2.0. Client validation completed April 2024; 128 GB version in development.
- Micron: CXL modules announced, enterprise positioning.
Sources: Tom’s Hardware — Samsung CMM-D 512 GB; ServeTheHome — SK Hynix CXL 2.0.
In plain terms: Type 1 = smart network card. Type 2 = GPU with shareable memory (theoretical for high-end AI). Type 3 = DRAM enclosure plugged into PCIe — the one that actually exists commercially in 2025.
LLM use case — KV cache for long-context inference
This is where CXL finds its strongest justification for AI workloads in 2025–2026.
The bottleneck: KV cache overflows available memory
When an LLM generates text on a long context, it must retain intermediate attention representations for each token already read. This structure is the KV cache — keys (K) and values (V) from each attention layer, for each token. It grows linearly with context length and layer count.
For a 70B model with 80 layers, in BF16 precision, on a 128,000-token context: the KV cache alone approaches 80 GB per request. On an H100 (80 GB total HBM), this leaves no room for the model weights (which are themselves spread across multiple GPUs). The current dominant solution, vLLM with PagedAttention (Kwon et al., SOSP 2023), manages the KV cache in pages within GPU VRAM. But VRAM remains the hard constraint.
The idea of KV cache offload to CXL: store the KV cache in an external DRAM pool (cheap, virtually unlimited capacity) and keep in HBM only the active tokens of the current context. The host CPU accesses the CXL memory via native load/store, without NIC, without network stack.
Advantage over RDMA: RDMA (direct memory access to a remote server via InfiniBand) requires explicit synchronization primitives, involves the network stack (NIC + driver + OS), and adds several microseconds of latency. CXL offers lower latency and access semantics identical to local DRAM.
Recent research results [early research]
Beluga (arXiv 2511.20172, accepted SIGMOD 2026, November 2025): first system enabling GPUs to directly access large memory pools via CXL switches. Announced results: 89.6% reduction in TTFT (Time-To-First-Token), 7.35x throughput gain in vLLM versus an RDMA baseline. [early research]
TraCT (arXiv 2512.18194, December 2025): rack-scale disaggregated LLM serving with KV cache shared via CXL. The system separates prefill (compute-intensive, GPU) and decode (latency-critical) phases, distributing them across separate nodes sharing their KV cache through a CXL pool. To manage coherence with non-coherent inter-node CXL, TraCT implements a two-tier synchronization mechanism. Announced results: TTFT reduced up to 9.8x, P99 latency improved 6.2x, peak throughput improved 1.6x versus RDMA/DRAM baselines. [early research]
NeurIPS 2024 ML for Systems Workshop (Tang et al.): exploration of KV cache on CXL storage. The paper documents that CXL provides native hardware support for load/store primitives, guaranteeing fine-grained low-latency access suited to the sparse access patterns of KV cache. [early research — workshop, not replicated]
Hardware demo XConn + MemVerge (Supercomputing 2025): demonstration of dynamic KV cache offload on CXL pool via XConn SC50256 switch. Announced result: over 5x gain versus SSD or RDMA. First commercial hardware demonstration at scale. [Source: Introl communiqué 2025, triangulated]
Microsoft Azure (November 2025): launch of first cloud instances equipped with CXL memory, marking the transition from research to real cloud deployment. Performance figures not published.
In plain terms: instead of limiting LLM context to what fits in GPU VRAM, CXL allows offloading the KV cache to an inexpensive external DRAM pool, accessible like local memory. Late-2025 research results show significant gains in latency and throughput — but these figures come from lab prototypes, not production deployments.
Limits and maturity — what benchmarks say
CXL latency is higher than local DRAM
This is the fundamental technical limitation of CXL, and it is precisely documented. The ASPLOS 2025 study (Jiang et al., “Dissecting CXL Memory Performance at Scale”, arXiv 2409.14317) tested 4 real CXL 1.1 devices across 265 workloads:
| Memory type | Typical latency |
|---|---|
| Local DDR5 DRAM | 81–114 ns |
| CXL 1.1 device (without NUMA) | 214–394 ns |
| CXL with NUMA crossing | 333–621 ns |
This represents a 2.6 to 4.9x latency increase over local DRAM. In bandwidth, tested devices delivered 18–52 GB/s, versus 52–246 GB/s for local DRAM depending on platform.
OSDI 2024 (“Managing Memory Tiers with CXL”) measures 2.02x load-to-use latency on Intel Xeon Gen 5 with a real CXL device, confirming the order of magnitude.
Impact on KV cache: the KV cache access pattern is sparse and often non-sequential (access to distant tokens in context). CXL latency is therefore more penalizing than for sequential access. However, CXL remains far faster than NVMe SSD (latency ~100 µs) and RDMA (latency ~2–5 µs but with protocol overhead). CXL sits between local DRAM and NVMe — an intermediate tier.
CXL 3.0 sharing requires software synchronization
CXL 3.0 promises hardware memory sharing (multiple hosts reading the same bytes), but inter-node coherence is not guaranteed without an additional software protocol. TraCT (arXiv 2512.18194) documents the complexity of a two-tier mechanism to ensure KV cache consistency across prefill and decode nodes.
Hyperscaler adoption vs. the rest of the market
| Segment | 2025–2026 situation |
|---|---|
| Hyperscalers (Microsoft Azure, targeted 2025–2026) | First CXL instances live (Microsoft, Nov. 2025). Progressive adoption on memory-intensive workloads. |
| Colocation / Edge | Near-zero deployment. Software stack (Linux drivers, tiered memory management) too complex without a dedicated team. |
| HPC / AI Research | Supercomputing 2025 demos (XConn + MemVerge). Mass adoption estimated 2027 after software stack maturation (ServeTheHome). |
ABI Research data cited by some sources (10% of servers 2024, 50% of AI racks 2025) could not be verified from their primary source [NOT VERIFIED].
In plain terms: CXL is 2.6 to 5x slower than local DRAM depending on generation and topology. This slowdown is acceptable for KV cache (workload tolerant to latency on distant tokens) but penalizing for bandwidth-bound workloads without alternatives. Microsoft Azure marks the first real cloud adoption in November 2025.
The SemiAnalysis controversy — “CXL is dead for GPU AI”
An argument has circulated in the industry since 2024: CXL would be structurally unsuited for high-end AI accelerators. The often-cited SemiAnalysis analysis argues that PCIe interfaces consume shoreline (the die edge surface available for external connections), which would be 3x less efficient in bandwidth per mm² than proprietary SerDes like NVLink or Google’s integrated interconnects. H100 and MI300X GPUs therefore allocate their shoreline to HBM links and NVLink — not to PCIe/CXL.
The essential nuance: this argument applies to GPU compute accelerators (Type 2) and the scale-up GPU-memory connection model. It does not apply to CPU servers with Type 3 memory expanders, where CXL is already deployed and useful for DRAM capacity expansion.
The CXL market is therefore structured around two distinct and non-competing segments:
- CPU memory expansion (Type 3): commercially viable since 2024, growing.
- GPU AI accelerators (Type 2): near-absent in high-performance production, replaced by proprietary HBM/NVLink.
The “CXL is dead” argument correctly describes the second segment. It does not describe the first.
Source: SemiAnalysis, “CXL is Dead in the AI Era” (paid newsletter, ~2024).
Decision matrix — when to use CXL
| Context | Recommendation | Why |
|---|---|---|
| CPU server with insufficient DRAM, memory-intensive workload, cloud or datacenter with a competent infra team | CXL Type 3 memory expander | Increases capacity beyond the physical DIMM slot ceiling without duplicating compute. 2–5x local DRAM latency acceptable for non-latency-critical workloads. |
| Long-context LLM KV cache (128K+ tokens), CPU server as orchestration tier | CXL pool (2.0+) for KV cache offload | Offloads KV cache from GPU HBM. Beluga/TraCT gains promising but [early research] — needs production validation. |
| Multiple CPU servers sharing a corpus in read-mode (e.g., embedding store) | CXL 2.0 pooling with dedicated partitions | Each host has its partition, switch resizable without reboot. Suited to sharing static data. |
| High-performance GPU (H100, MI300X) — scale-up AI acceleration | NVLink + HBM | CXL cannot compete with intra-GPU bandwidth. Die shoreline better used by HBM/NVLink. |
| Edge / colocation without dedicated infra team | Local DRAM + NVMe SSD | CXL software stack (drivers, OS tiered memory) too complex in 2025–2026 outside hyperscalers. |
Heuristic rules:
- Start with local DRAM: CXL is an additional tier, not a replacement. Fill DIMM slots before considering CXL expansion.
- Measure latency sensitivity: if the workload is bandwidth-bound with dense sequential access, the 2–5x CXL latency factor translates directly into slowdown. If accesses are sparse or latency can be masked (prefetch, async), CXL is viable.
- Wait until 2027 for edge production: the software stack (Linux CXL tiered memory drivers, NUMA management with CXL) will reach production maturity after 2026 per industry estimates.
- Distinguish pooling from sharing: pooling (CXL 2.0 — dedicated partitions) is stable and deployed. True sharing (CXL 3.0 — same bytes across hosts) requires additional software synchronization and remains [early research] for LLM workloads.