In brief

Training a large language model is a schedulable, batch operation spanning several weeks. Serving it in production is a radically different problem: every response must arrive quickly, thousands of users may connect simultaneously, and GPU memory is a scarce and expensive resource. Five families of techniques have emerged to address this challenge — each attacking a specific bottleneck.

In short: training a model is building a car at a factory; serving the model is running a delivery service where a thousand customers order at the same time. The optimization techniques all serve to unblock a specific point of the service — trunk space (memory), loading speed (KV cache), the number of simultaneous trips (batching), or a shortcut to guess the order (speculative decoding).


The problem: why inference is expensive

A 70-billion-parameter model stored in FP16 precision occupies approximately 140 GB of GPU memory. A single high-end A100 GPU has 80 GB. Already, the model does not fit.

But memory is not the only problem. Text generation is sequential by nature: each new token depends on all preceding tokens. The model must therefore run its full computation to produce a single word. And to avoid recomputing the same things at each step, systems keep intermediate attention results in memory (the KV cache — for Key-Value cache). This cache grows with each generated token, and even faster when multiple requests are processed in parallel.

The inference industry is estimated at $106 billion in 2025, projected to exceed $250 billion by 2030. Reducing this cost is as much a commercial challenge as a technical one.


Quantization: fitting into less memory

Quantization reduces the numerical precision of a model’s weights. Instead of storing each parameter on 16 bits (FP16), the precision drops to 8, 4, or even 2 bits. The gain is immediate: a model in 4-bit precision occupies four times less memory than in FP16.

The challenge is compressing without degrading performance. Three approaches have shaped the state of the art:

GPTQ (Frantar et al., ICLR 2023) uses second-order mathematical information to quantize layer by layer with minimal loss. On a 175-billion-parameter model, it enables approximately 3× speedup on an A100 GPU in 4 bits, while fitting the model into a single GPU.

AWQ (Lin et al., MLSys 2024, Best Paper Award) starts from an observation: not all weights contribute equally to the output. By identifying the 1% of most important channels (identified via activations, not weights), AWQ protects them from aggressive compression. This selectivity enables more than 3× speedup even on mobile GPUs.

SmoothQuant (Xiao et al., 2023) addresses a harder case: quantizing both weights and activations to INT8. Activations pose a specific problem — a few dimensions contain extreme values that resist compression. SmoothQuant mathematically shifts this difficulty onto the weights, which are easier to compress. Result: up to 1.56× speedup and 2× memory reduction.

An important nuance: 4-bit quantization is often presented as near-lossless on standard benchmarks. Recent work nuances this conclusion — mathematical reasoning and programming tasks suffer clear degradation under INT4, invisible on standard perplexity measures (arxiv 2505.11574). 4-bit compression preserves reasoning capabilities less well than perplexity suggests.

For comprehensive coverage of quantization methods (GGUF, llama.cpp, NF4, recent hardware formats), see the dedicated article.


FlashAttention: reducing round-trips inside the GPU

The attention operation — the central mechanism of Transformers — is known for its quadratic memory consumption with sequence length. But Tri Dao and co-authors (NeurIPS 2022) identified a more fundamental problem: the bottleneck is not the number of computations, it is the time spent moving data between the GPU’s high-bandwidth memory (HBM) and its fast on-chip memory (SRAM).

Standard attention reads and writes to HBM at every step. FlashAttention breaks computation into blocks so that it stays entirely within the much faster SRAM. Result: memory transfers go from quadratic to linear. Measured gain: 3× on GPT-2, without any modification to the output — the algorithm is exact, not approximate.

FlashAttention-2 (ICLR 2024) improves internal parallelism and achieves 50–73% of the theoretical maximum computation of an A100 GPU, compared to 25–40% for the previous version. FlashAttention-3 (2024) adds optimizations specific to Nvidia’s H100 GPUs.

FlashAttention is now integrated by default in most inference frameworks. Its impact is not marginal: it is one of the rare optimizations that simultaneously accelerates both training and inference.


vLLM and PagedAttention: managing the KV cache like an operating system

The KV cache is unpredictable: its size depends on the length of each generated response, which varies from one request to the next. Naive implementations reserve memory for the maximum possible size — which wastes 60 to 80% of GPU resources according to the vLLM team’s measurements.

Kwon et al. (SOSP 2023) borrowed an idea from operating systems: virtual memory and paging. In PagedAttention, the KV cache is broken into fixed-size blocks allocated on demand from a shared pool, exactly as an OS manages its memory pages. This eliminates fragmentation and also allows sharing common blocks between similar requests (for example, multiple requests using the same prefix).

vLLM builds on PagedAttention a complete serving system, with a key feature: continuous batching. Instead of waiting for all requests in a batch to complete before processing new ones, new requests join the batch as soon as slots become available. The measured gain is 2 to 4× throughput improvement over previous systems at equivalent latency.

An architectural debate is ongoing: vAttention (arxiv 2405.04437) proposes managing KV cache memory directly via operating system primitives, without vLLM’s paging mechanism, arguing that modern abstractions make PagedAttention unnecessarily complex. This debate remains open.


Speculative decoding: parallelizing a sequential process

The fundamental constraint of autoregressive generation is simple: to produce token N, token N-1 must already have been produced, meaning the large model must run at least once per token. Parallelizing this process seems impossible — in appearance.

Speculative decoding (Leviathan et al., ICML 2023; Chen et al., 2023) proposes a workaround: a smaller auxiliary model (the draft model) proposes K tokens ahead, then the large model verifies all of them in a single parallel pass. If the proposals are accepted, K tokens are produced at the cost of a single call to the large model.

The mathematical guarantee is strong: the procedure exactly preserves the probability distribution of the large model. The result is identical to what it would have produced without acceleration.

Modern implementations have abandoned the external draft model:

Medusa (Cai et al., 2024) adds extra decoding heads directly to the model, predicting multiple next tokens in parallel. Gain: 2.2 to 3.6× depending on configuration.

EAGLE-3 (Li et al., 2025) uses a one-layer mini-Transformer as a draft, reusing the large model’s internal representations. Gains reach 3 to 6.5× versus standard generation.

A documented limitation: the gains from speculative decoding strongly depend on the acceptance rate of proposed tokens, which varies across languages and tasks. Recent work (arxiv 2510.02128) shows that languages underrepresented in training data have much lower acceptance rates — creating a latency inequality between users depending on their language.


Distillation: building smaller but capable models

Knowledge distillation is a different approach: rather than optimizing the deployment of a large model, a smaller model is trained to reproduce the large model’s behavior. The student does not learn from the original data, but from the teacher model’s outputs — its probability distributions, its intermediate reasoning.

This approach takes two main forms in the LLM context:

  • Output distillation: the small model learns from texts generated by the large model. This is the dominant paradigm since GPT-4: many open-source models have been fine-tuned on synthetic data produced by proprietary models.

  • Intermediate representation distillation: the student is supervised on the large model’s internal activations, not just its outputs. More complex to implement, but more precise.

Chain-of-Thought distillation forces the student to reproduce the intermediate reasoning steps, not just the final answer. This technique partly explains why 7- or 14-billion-parameter models can compete on certain tasks with models ten times larger.

An open question: training an open-source model on GPT-4 or Gemini outputs violates those services’ terms of use. The legal status of this practice is unresolved and weighs on the reproducibility of a significant portion of published distillation work (Xu et al., arxiv 2402.13116).

For comprehensive coverage of distillation methods (families, pruning, self-distillation, limitations), see the dedicated article.


Five bottlenecks, five responses

These techniques are not redundant — they address different constraints:

TechniqueTargeted bottleneckMechanism
QuantizationGPU memoryReducing weight precision
FlashAttentionInternal memory bandwidthBlock computation in SRAM
vLLM / PagedAttentionKV cache fragmentationOn-demand paging
Speculative decodingSequential generationDraft + parallel verification
DistillationModel sizeTransferring capabilities to a smaller model

In practice, they are combined. A typical production system uses a model quantized to 4 bits, served via vLLM with FlashAttention-2, with active speculative decoding. The challenge is that these optimizations interact: 4-bit quantization can reduce the quality of the draft model’s proposals, and continuous batching can conflict with static quantization assumptions. Globally optimizing an inference system remains an open problem.


Which technique for which bottleneck

ContextPriority recommendationWhy
Model does not fit in memory (single GPU)INT8 or INT4 quantizationCuts size by 2 to 4×, degradation below 1 point on standard benchmarks from INT8 up.
Time to first token (TTFT) too longFlashAttention v2/v3Reduces HBM ↔ SRAM memory traffic by 5-10× on long prompts. Enabled by default in vLLM and SGLang.
Insufficient throughput (requests per second plateauing)vLLM + PagedAttentionReduces KV cache fragmentation, multiplies concurrent requests by 2-4×.
Inter-token latency (TPOT) too longSpeculative decoding (draft model + verification)2-3× faster on long generations, at negligible cost to quality.
Marginal cost too high (production scale)Distillation to a 7B-13B model5-10× cheaper to serve than a 70B+, task-dependent degradation (often under 5 points on calibrated tasks).
Power and carbon cost problematicQuantization + distillation + edge inference combinedMoves part of the load to CPUs or embedded GPUs, taking the datacenter off the hot path.

Key takeaways

  • A 70-billion-parameter model at standard precision exceeds the capacity of a single high-end GPU — quantization (GPTQ, AWQ) allows it to fit in 4 bits with quality loss that is often minor, but measurable on reasoning tasks.
  • FlashAttention solves an internal GPU memory transfer problem, not a computation problem — which is why its gains compound with other optimizations.
  • vLLM introduces OS-inspired paging for KV cache management and multiplies throughput by 2 to 4× compared to naive systems.
  • Speculative decoding bypasses the sequentiality of generation using a fast draft model, with a mathematical guarantee of preserving the large model’s distribution — but its gains are uneven across languages.
  • Distillation produces smaller models capable of behavior close to larger models, but raises unresolved legal questions when it relies on the outputs of proprietary models.