In brief

A 70-billion-parameter model stored in FP16 occupies roughly 140 GB of GPU memory. No consumer PC can host it. Quantization reduces the numerical precision of each weight — from 16 bits to 4 bits, for example — bringing that same model down to 35 GB. Quality drops slightly. But usually less than feared.

In short: quantization is the mover who reduces the size of each box. You don’t throw out any objects, you just pack them more tightly. In the 16-bit box, each object has its own precise compartment. In the 4-bit box, several objects share the same compartment — you accept slightly blurring the nuances. The move costs fewer trucks (memory), goes faster (latency), but you sometimes find two pairs of socks mixed up.


The starting problem: floating-point numbers are expensive

A neural network is a massive list of numbers — the weights. Each weight encodes a fraction of what the model has learned. During training, these numbers are stored in FP32 (32-bit floating point) for maximum precision. At inference, we often switch to FP16 or BF16 (16 bits): a 2× memory gain with negligible quality loss.

But 16 bits per weight, over 70 billion weights, amounts to 140 GB. To fit on a consumer GPU (16–24 GB of VRAM), you need to go further.

Quantization pushes this logic further: represent each weight with 8, 4, or even 2 bits. Memory drops. Inference speeds up. The cost: an information loss that the model must absorb without degrading too much.

In short: going from FP16 to INT4 means dividing memory size by 4. A 70B model goes from 140 GB (impossible on a consumer GPU) to 35 GB (fits on an RTX 4090 + a bit of DRAM). You access models that otherwise could only be served from a datacenter. Quality loss is measurable but limited — typically 1-3 benchmark points on INT4, much more on INT2.

The precision scale

The standard progression of formats runs from high precision to extreme compression:

  • FP32 (32 bits): training, maximum precision
  • FP16 / BF16 (16 bits): standard inference, near-zero loss
  • INT8 / FP8 (8 bits): 2× additional compression, <2% perplexity degradation
  • INT4 / NF4 (4 bits): 4× vs FP16, 2–8% degradation
  • 3 bits: 8–15% degradation, experimental use
  • 2 bits: 15–30% degradation, research territory

Each step roughly halves memory. Degradation is not linear: going from 16 to 8 bits is nearly free; going from 4 to 2 bits starts to hurt.

Two families of methods

PTQ — Post-Training Quantization

Quantization is applied after training, on an already-trained model. A small calibration dataset (a few hundred examples) is used to estimate how to compress the weights with minimal damage. Cost: a few hours of compute, no A100 GPU required. This is the dominant method in practice.

QAT — Quantization-Aware Training

Quantization is simulated during training or fine-tuning. The model learns to operate with compressed weights. Result: better quality, especially below 4 bits. Cost: prohibitive for large models. On smaller models, PyTorch QAT recovers up to 96% of the degradation observed in PTQ on the HellaSwag benchmark with Llama 3. For a 70B model, no viable QAT pipeline at this scale has yet been published.

The PTQ methods that matter

GPTQ

GPTQ (2022–2024) was the first method to make 4-bit quantization genuinely viable. It works layer by layer: for each transformer layer, it quantizes weights by compensating for the error introduced on subsequent weights, via an algebraic reformulation based on the Hessian (Optimal Brain Quantization). It is the most widely used approach on HuggingFace.

AWQ

AWQ (Activation-Aware Weight Quantization) starts from an observation: not all weights are equal. Roughly 1% of channels are “salient” — they disproportionately influence the output, as revealed by activation analysis. AWQ identifies these channels and applies a scaling before uniform quantization, preserving the essential information. Result: AWQ regularly outperforms GPTQ in precision at equivalent compression ratios.

SmoothQuant

SmoothQuant targets a different objective: quantizing both weights and activations (W8A8 format). The challenge with activations is the presence of outlier values that are difficult to compress. SmoothQuant mathematically transfers the difficulty from activations to weights, making both uniformly quantizable. This approach is particularly suited to hardware accelerators optimized for W8A8.

AQLM and SpQR — for extreme compression

To go below 3 bits, two methods stand out. SpQR combines sparse representation and quantization: the most important weights stay in high precision, the others are aggressively compressed, with perplexity loss below 1% at roughly 4 bits per parameter. AQLM pushes down to 2 bits via additive per-block quantization with joint codebook optimization — the first Pareto-optimal scheme below 3 bits according to the literature.

Concrete formats

INT8 and FP8: the safe zone

INT8 is the most stable and widely supported format. Quality loss is negligible (<2% perplexity). FP8 (8-bit floating-point format, variants E4M3 and E5M2) is even more stable on large models: it is the recommended option for production deployment according to an IJCAI 2025 study covering multiple architectures and tasks. NVIDIA has integrated FP8 natively in its Hopper GPUs (H100); vLLM 0.5 (2024) achieves a ~33% latency gain from it.

NF4 (NormalFloat 4 bits) is a 4-bit floating-point format designed for normally distributed weights — which describes the majority of LLM weights. Introduced by the QLoRA paper (Dettmers et al., 2023), it is used via the bitsandbytes library for fine-tuning on consumer GPUs. Degradation stays in the 2–8% range depending on model and task.

GGUF and llama.cpp: the CPU ecosystem

GGUF is the standard file format for local CPU inference, developed by the llama.cpp project. It replaces the older GGML format and supports compression levels from Q2 to Q8, with K-quant variants (suffixes _S, _M, _L for small/medium/large).

The recommended trade-off is Q4_K_M: a 7B model goes from 13.5 GB (FP16) to 4.08 GB, retaining roughly 95% of quality. For critical use cases — coding, STEM — Q5_K_M is preferable. llama.cpp enables native CPU inference without a GPU, explaining its massive adoption for local deployment.

NVFP4: the format for next-generation datacenters

Launched in 2025 for Blackwell GPUs (B200), NVFP4 is a 4-bit floating-point format with a block size of 16. Compared to FP16, it reduces memory by 3.5×. Compared to FP8, it offers 2.3× more throughput with less than 1% degradation on language tasks. It is the target format for next-generation high-performance deployments.

What resists compression

Models do not age equally well under compression. Some empirical observations from the literature:

Coding and STEM tasks suffer more. General text generation is more robust than exact code production or mathematical reasoning. A model quantized to 4 bits on a conversational generation task may appear intact while producing more errors on code.

Small models (<7B) are more fragile. Low-precision quantization is calibrated for large models. Below 7B parameters, losses are disproportionate. SLMQuant (November 2025) documented this phenomenon without proposing a general solution.

A little-known risk: package hallucinations. A 2025 study shows that 4-bit quantization significantly increases the rate of software package name hallucination in code generation tasks — quantized models invent more nonexistent or vulnerable libraries. This risk is not covered by standard perplexity benchmarks.

The tooling

To deploy a quantized model, several tools are essential depending on use case:

  • llama.cpp — CPU/GPU inference in C++, GGUF format, cross-platform. De facto standard for local deployment.
  • bitsandbytes — PyTorch library for loading models in NF4/INT8 on the fly. Essential for QLoRA and fine-tuning on consumer GPUs.
  • vLLM — high-performance GPU serving, supports FP8, GPTQ, AWQ, NVFP4. Direct integration with LLM Compressor (Red Hat).
  • TensorRT-LLM (NVIDIA) — datacenter deployment optimization, FP8/NVFP4 on Hopper and Blackwell. Open-source since January 2025.
  • ExecuTorch (Meta) — edge runtime, general availability October 2025, 50 KB base, supports over 80% of edge LLMs available on HuggingFace.

Open questions

PTQ vs. QAT at scale. QAT is superior in quality but its cost remains prohibitive beyond 70B parameters. A hybrid PTQ→QAT pipeline is emerging but not yet standardized.

Quantization of MoE models. Since early 2025, more than 60% of frontier models use a Mixture of Experts architecture. Quantizing these models raises specific questions — active vs. inactive experts — that are poorly documented in the literature.

Security impact. The effect of quantization on alignment properties, RLHF guardrails, and jailbreak resistance is almost unexplored. Only one recent paper (arXiv:2502.15799) addresses the topic; its conclusions have not yet been replicated.

Evaluation metrics. Perplexity is the dominant but insufficient proxy. No unified post-quantization benchmark exists covering coding, STEM, security, and hallucination. This is a structural blind spot in published comparisons.


Which format for which hardware and use

ContextRecommendationWhy
Datacenter, recent GPU (H100, B200), productionFP8 or INT8Quality nearly identical to FP16, 2× memory gain, latency -30 to -50%. NVFP4 on Blackwell pushes further still.
Workstation, 24 GB consumer GPU (RTX 4090)GPTQ INT4 or AWQ INT4Lets a 30-70B model fit in VRAM. Degradation of 1-3 benchmark points. Mature libraries (vLLM, ExLlamaV2).
Laptop, CPU + integrated GPUGGUF Q4_K_M or Q5_K_M (llama.cpp)Format optimised for CPU/Apple Silicon. Q5_K_M is the best quality/size balance in most cases.
Edge / mobile, hard memory constraintAQLM or SpQR (extreme INT2-3 compression)Makes serving 7B models on high-end smartphones possible. Stronger degradation (5-10 points), to be validated task by task.
Research, sensitivity to outliers (maths, code)SmoothQuant or QATPreserves the large-amplitude activations that derail naive methods. More expensive to apply.

Key takeaways

  • Quantization reduces weight precision (FP16 → INT8 → INT4) to decrease memory and accelerate inference.
  • INT8 and FP8 offer 2× compression with negligible degradation. INT4 offers 4× with moderate but acceptable degradation in most use cases.
  • PTQ (post-training) dominates in practice; QAT is superior but reserved for small models or teams with fine-tuning resources.
  • GPTQ and AWQ are the reference PTQ methods; AWQ is generally more precise. GGUF/llama.cpp is the standard for local CPU inference.
  • Models <7B, coding tasks, and code generation are more sensitive to degradation than standard benchmarks suggest.
  • The impact on security and alignment remains a grey zone: understudied, potentially non-negligible.