In short
The Transformer is the architecture behind nearly every large language model in use today — GPT, Claude, LLaMA, Gemini, Mistral. Introduced in 2017 by Google researchers, it replaced recurrent networks by putting a single principle at the center: attention. Understanding that mechanism is the key to understanding why these models can hold a coherent thread over long texts, translate, code, summarize — and also why they have very specific limits.
In short: a Transformer is like a classroom where every student (every word in the text) can spontaneously raise their hand and look at every other student at once to understand their own place in the topic. Before the Transformer, students had to pass a note in relay (recurrent model) — which works for five students, breaks down at a hundred. With the Transformer, everyone sees each other at a glance. The cost: you need a room large enough for everyone’s cross-glance (memory in N²).
The starting intuition: put attention at the center
Before 2017, language-processing models read sentences word by word, in sequence. A recurrent network read “The cat sleeps on the rug” from left to right, storing a compressed memory at every step. That compressed memory quickly fell short when it came to capturing relationships between words far apart in the text.
The Transformer solves this differently: instead of reading sequentially, it looks at every word at once and, for each word, computes how much it should “pay attention” to each of the others. This is the principle of self-attention.
A useful analogy: think of a document lookup. To answer a question (the query), you browse a library (the keys) and pull out the content of the relevant books (the values). The attention mechanism does exactly that, but for every token in a sentence, against every other token.
Formally, for a given sequence, three matrices are computed — Q (queries), K (keys), V (values) — and attention is written as:
Attention(Q, K, V) = softmax(QKᵀ / √d_k) · V
The √d_k factor stabilizes the calculation when the representation dimension is large. The result: a weighting over every position in the sequence, for every position. A model can therefore directly connect “he” and “Napoleon” in “Napoleon lost the battle because he was sick”, regardless of the distance between the two words.
In short: self-attention is the word “he” scanning every other word in the sentence and instantly choosing which one it refers to — Napoleon, rather than the battle or the illness. No relay, no compressed memory to reconstruct. Every word sees everything. It is this ability to directly connect two distant words that gives Transformers their coherence over long texts.
Multi-head attention: several simultaneous readings
A single attention layer is not enough. The Transformer uses multi-head attention: several attention “heads” operate in parallel, each with its own projections. Some heads capture syntactic relations (subject–verb), others semantic ones (synonyms, anaphora), others long-range dependencies.
The outputs of all heads are concatenated, then re-projected. It is this richness of simultaneous viewpoints that gives the Transformer its fine-grained understanding.
Positional information: a problem to solve
The attention mechanism is, by construction, indifferent to token order. “Cat eats mouse” and “Mouse eats cat” would produce the same representations without correction. Position information must therefore be injected explicitly for every token.
Three approaches dominate:
Sinusoidal encoding (2017): vectors based on sine and cosine functions are added to the input representations. Simple, but models extrapolate poorly to sequences longer than those seen during training.
RoPE — Rotary Position Embedding (2021): position is encoded as a rotation in the vector space. The rotation is absolute per token, but the comparison between two tokens retains only their relative position. RoPE has become the standard for modern LLMs (LLaMA, Mistral, DeepSeek, Qwen). Its limit: extrapolating beyond the training window remains hard, which motivated extensions like YaRN (2023) and LongRoPE.
ALiBi — Attention with Linear Biases (2021): instead of adding vectors, ALiBi directly penalizes attention scores based on the distance between tokens. The farther apart two tokens are, the more their score is penalized. No position vector is added to the representations. ALiBi generalizes better to long sequences. Adopted in MPT and BLOOM.
A comparative study published at EMNLP 2024 tempers the debate: after fine-tuning, models trained with these three methods tend to converge toward similar performance.
Three architectural families
The original Transformer architecture was encoder-decoder. Since then, three families have emerged depending on use case:
| Architecture | Attention | Examples | Main use |
|---|---|---|---|
| Encoder-only | Bidirectional (sees everything) | BERT, RoBERTa | Understanding, classification, NER |
| Decoder-only | Causal (only sees the past) | GPT, LLaMA, Claude, Mistral | Text generation, LLM |
| Encoder-decoder | Bidirectional encoder + causal decoder | T5, BART | Translation, summarization, generative QA |
The current dominance of decoder-only is not a universal technical given. Results published in 2025 suggest encoder-decoder can outperform decoder-only after instruction-tuning at equal compute budget. The question remains open: the popularity of decoder-only may reflect the historical success of GPT as much as an intrinsic architectural superiority.
Scaling laws: how many parameters, how much data?
A fundamental question for LLMs is scaling: how does performance evolve as you scale up model size, data volume, or compute budget?
Kaplan et al. (2020, OpenAI) showed that model loss follows stable power laws across seven orders of magnitude. Their recommendation at the time: at a fixed compute budget, invest most of it in model parameters.
Hoffmann et al. (2022, DeepMind) revisited that conclusion with Chinchilla: by training more than 400 models from 70M to 16B parameters, they showed that Kaplan underestimated the importance of data. The revised rule: model size and training-token count should grow proportionally. Chinchilla (70B parameters, 1.4 trillion tokens) thus outperformed Gopher (280B) and GPT-3 (175B) on many benchmarks.
But recent models — LLaMA 3, Mistral — are deliberately trained beyond Chinchilla’s recommendations. The reason: a smaller but over-trained model is cheaper to deploy at scale. Scaling laws now incorporate inference cost, not just training cost.
What changes at scale: efficiency and trade-offs
Two practical problems shape the engineering of modern Transformers.
The quadratic complexity of attention: computing attention between every pair of tokens costs O(n²) in memory and compute. For a 100,000-token sequence, this quickly becomes prohibitive. FlashAttention (Dao et al., 2022–2024) partially solves the problem: by reorganizing memory accesses to exploit the GPU hierarchy, required memory drops to O(n) with a 2–4× speedup. Time complexity stays quadratic, but the problem becomes manageable in practice.
The KV-cache: at inference time (when the model generates text), the keys and values of every token already produced are cached to avoid recomputing them. This cache consumes a lot of memory. Two attention variants reduce that cost:
- Multi-Query Attention (MQA): a single K/V head shared across all query heads — faster, slight quality loss.
- Grouped Query Attention (GQA): a few groups of heads each share a K/V head — a trade-off between speed and quality, adopted in LLaMA 2 (70B), Mistral, Gemma.
What researchers are still debating
The Transformer architecture has stabilized, but several questions remain open.
Are emergent capabilities real? In 2022, several teams observed that certain capabilities (arithmetic reasoning, analogical understanding) seemed to appear abruptly at certain size thresholds. In 2023, Schaeffer et al. published a strong counter-argument: these “jumps” are artifacts of discontinuous metrics. With continuous metrics, progress would be smooth. The debate is not settled.
Can Transformers be replaced? Alternative architectures — Mamba, RWKV, Hyena — rely on sub-quadratic mechanisms that could cut the cost of long sequences. In practice, they still struggle to match Transformers on standard benchmarks. The most promising direction seems to be hybrid architectures combining local attention with alternative mechanisms.
Length generalization remains an unsolved problem. Even with RoPE and its extensions, models struggle to reason reliably over contexts well beyond their training window. YaRN, LongRoPE, Position Interpolation are empirical fixes, without solid theoretical grounding.
Which architectural family for which need
| Context | Recommendation | Why |
|---|---|---|
| Fine-grained text understanding (classification, NER, sentiment) | Encoder-only (BERT, RoBERTa, DeBERTa) | Optimized to understand the full text in one pass. No generation but low inference cost. |
| Conversational or creative text generation | Decoder-only (GPT, Claude, LLaMA, Mistral) | Dominant family 2024-2026. Generates token by token, capable of zero-shot and instruction-following. |
| Translation, summarization, text→text transformation | Encoder-decoder (T5, BART, Flan-T5) | Classical seq2seq model. Rarer today, but still relevant for supervised translation tasks. |
| Long context (> 100K tokens) | Decoder-only with attention optimization (Mistral, Claude, Gemini, Mamba hybrid) | FlashAttention, sparse attention, Mixture-of-Experts or state-space hybrid to support 200K-1M tokens. |
| Low latency, edge inference | Small quantized decoders (Llama 3.2 1B-3B, Phi-3 mini) | Workable on CPU/mobile. Limited performance but usable outside the datacenter. |
What to remember
- The Transformer is built on attention: every token simultaneously computes its relationship with every other token, with no sequential processing.
- Multi-head attention allows multiple types of relationships to be learned in parallel — syntax, semantics, long-range dependencies.
- Positional encoding (RoPE dominant today) is a necessary addition because attention is indifferent to token order.
- Scaling laws (Kaplan 2020, Chinchilla 2022) guide model sizing; recent practice goes beyond Chinchilla by factoring in inference cost.
- The dominance of decoder-only (GPT, LLaMA, Mistral) is as much historical as technical — the debate with encoder-decoder remains open.
- The big open questions: quadratic complexity at very long sequences, generalization beyond the training window, reality of emergent capabilities.