In brief
Current large language models are the result of five successive breakthroughs, each unlocking the next: vector word representations (2013), the Transformer architecture (2017), the pre-training/fine-tuning paradigm (2018), the revelation of scale (2020), then alignment as a central challenge (2022). The popularity of ChatGPT in November 2022 is not a technical breakthrough — the underlying model had existed for months. It is an interface breakthrough: a powerful LLM made accessible to non-technical users for the first time.
In short: think of LLM history like aviation history. The Wright brothers (Word2Vec, 2013) proved that an object heavier than air could fly. The Transformer (2017) was the jet engine that multiplied the range. Pre-training (2018) made it possible to produce reusable cells. Scaling (2020) showed that a bigger plane carried farther. Alignment (2022) made the plane mainstream — it had to do what the passengers asked, not just what it could do.
2013 — Words as points in space
Before 2013, natural language processing systems treated words as discrete symbols, with no relationship between them. “King” and “queen” had no connection in that space; “bank” referred to the same thing whether discussing money or a river.
Tomas Mikolov and his colleagues at Google changed this with Word2Vec (2013). The idea: represent each word not as an isolated symbol, but as a point in a geometric space of several hundred dimensions. A word’s position in this space captures its relationships with other words. The now-famous demonstration: in this space, king − man + woman ≈ queen. Vector arithmetic preserves semantic structure.
This approach is called word embedding. Stanford refined the method in 2014 with GloVe, which combines global co-occurrence statistics with local learning. These vector representations formed the backbone of natural language processing until 2018.
Their limitation: a word receives the same vector regardless of context. “Bank” always has the same representation, whether the sentence is about finance or geography. ELMo (2018) would be needed to overcome this constraint.
2017 — The Transformer: when all words speak simultaneously
On June 12, 2017, a team from Google Brain published “Attention Is All You Need” (Vaswani et al., arXiv:1706.03762). The paper proposed a radically different architecture from existing networks — the Transformer — which abandons recurrent networks and their sequential processing.
The central mechanism is attention: each word in a sentence can directly “look at” all other words and weight their importance based on context. To understand “he” in “Paul saw John, he was tired,” a human connects “he” to Paul or John based on the overall meaning. The attention mechanism does the same — and does so simultaneously for all words in the sequence.
Two major consequences follow. First, parallelization: unlike recurrent networks that processed sequences step by step, the Transformer processes the entire input in a single pass, exploitable on large-scale GPUs. Second, capturing long-range dependencies: words far apart in a text can directly influence each other, without the memory issues affecting previous architectures.
On machine translation, the Transformer surpassed all previous models in less than a week of training on 8 GPUs. The breakthrough was undeniable.
In short: before 2017, understanding a sentence meant reading it word by word — like rewinding a cassette tape. With attention, the machine sees the entire sentence at once and instantly recognizes that “he” refers to “Paul” (and not “John”). This parallelization is not just an optimization: it is what made training on hundreds of billions of words possible, which was unthinkable with sequential architectures.
2018 — Pre-train, then fine-tune
2018 consolidated a paradigm that would shape the field: train a model on a very large unlabeled generalist corpus, then specialize it on specific tasks with far less data.
ELMo (Peters et al., NAACL 2018) marked the first break from static embeddings: using bidirectional networks, it produces contextual representations — the same word receives a different representation depending on its sentence. Added to existing models without modification, ELMo improved the state of the art on six natural language benchmarks.
OpenAI published GPT-1 the same year: a Transformer trained on 800 million words with a simple objective — predict the next word. Then fine-tuned on specific tasks. GPT-1 demonstrated that massive pre-training transfers effectively to highly varied tasks.
Google responded with BERT (Devlin et al., arXiv:1810.04805): a bidirectional encoder architecture that reads the sentence in both directions and predicts masked words instead of the next word. BERT set a new record on eleven language understanding tasks. The GPT (generative, left-to-right) versus BERT (encoder, bidirectional) opposition would structure architectural debates in the years that followed.
2020 — The revelation of scale
GPT-2 (2019, 1.5 billion parameters) had shown something surprising: the model accomplished translation or summarization tasks without having been explicitly trained for them, simply through conditioning in the prompt. OpenAI even delayed its full release citing misinformation risks.
GPT-3 (Brown et al., arXiv:2005.14165, 2020) takes this reasoning to its conclusion: 175 billion parameters, trained on approximately 500 billion words. The revealed phenomenon is called in-context learning: without any subsequent fine-tuning, GPT-3 receives a few examples in its context and generalizes immediately. This capability was not present in smaller models — it emerges with scale.
Kaplan et al. (arXiv:2001.08361, 2020) formalized an empirical observation: the quality of LLMs follows power laws as a function of the number of parameters, data size, and compute budget — across seven orders of magnitude. These scaling laws imply that larger mechanically means better, and oriented the race toward ever-larger models that characterized 2020–2022.
In 2022, a decisive correction came from DeepMind: the Chinchilla law (Hoffmann et al., arXiv:2203.15556). Existing models were massively undertrained on data relative to their size. Chinchilla, with 70 billion parameters but four times more data than its contemporaries, surpassed GPT-3 (175 billion) on nearly all benchmarks. The race for raw scale was tempered in favor of a better parameter/data balance.
In short: Chinchilla is the culinary equivalent of realizing that a dish does not need 100 ingredients to be good — it needs the right balance. With 70 billion well-fed parameters (1.4 trillion tokens), Chinchilla beats GPT-3 (175 billion but starved at 300 billion tokens). The empirical rule that emerges: 20 tokens per parameter is a better ratio than the simple race for model size.
2022 — Alignment: following intentions, not distributions
GPT-3 impressed researchers but disappointed users. The model followed statistical distributions, not human intentions: toxic responses, hallucinations, unpredictable behavior when given instructions.
The technical response is called RLHF (Reinforcement Learning from Human Feedback). The pipeline, formalized by Ouyang et al. from OpenAI (arXiv:2203.02155, 2022), works in three steps: supervised fine-tuning on human demonstrations, training a reward model on human comparisons, then reinforcement learning optimization to maximize that reward.
The result was striking: InstructGPT with 1.3 billion parameters was preferred by human annotators over GPT-3 with 175 billion in 85% of cases. A hundred times fewer parameters, the same level of preference. Size yielded priority to alignment.
ChatGPT (November 2022) applied this pipeline to a conversational interface. In five days, one million users. In two months, one hundred million — no technology product had reached this adoption so rapidly. The breakthrough was not in the model itself, but in access: a powerful LLM made usable without technical expertise.
Anthropic simultaneously published an alternative to RLHF: Constitutional AI (Bai et al., arXiv:2212.08073). Instead of human annotators, the model self-critiques according to an explicit set of principles and produces its own training data. The approach reduces dependence on costly human annotations.
In short: ChatGPT reached 100 million users in two months — a record surpassing TikTok (9 months) and Instagram (2.5 years). Yet the underlying model (GPT-3.5) had existed since the start of the year. It was not the technology that flipped in November 2022, it was the interface: going from a technical API to a chat box where you type in natural language. The lesson holds far beyond LLMs: a usage breakthrough can matter more than an algorithmic one.
2023–2025 — Fragmentation and reasoning
2023 saw the proliferation of frontier models. GPT-4 (OpenAI, 2023) integrated image understanding and surpassed GPT-3.5 on nearly all professional benchmarks — bar exams, medicine, mathematics. Gemini (Google DeepMind, 2023) is natively multimodal, trained simultaneously on text, images, audio, and video.
On the open-weights side, Meta published LLaMA (arXiv:2302.13971): models ranging from 7 to 65 billion parameters trained on entirely public data. LLaMA-13B surpassed GPT-3 (175B) on most benchmarks. The open-source community exploded. Mistral AI, founded in Paris in 2023, published Mistral-7B which surpassed LLaMA-2-13B despite having half the parameters, thanks to more efficient attention techniques.
In 2024–2025, a new breakthrough emerged around reasoning: rather than optimizing direct response, models are trained to produce long chains of thought before concluding. OpenAI published the o1/o3 series. DeepSeek (China) published DeepSeek-R1 (arXiv:2501.12948, January 2025) using pure reinforcement learning: the model spontaneously develops self-verification behaviors and exploration of alternatives. R1 rivals o1 at a fraction of the development cost, demonstrating that training engineering can compensate for resource gaps.
Open debates
Two questions remain without a definitive answer.
Are emergent capabilities real? Wei et al. (arXiv:2206.07682, 2022) describe capabilities absent in smaller models and suddenly present at scale. Schaeffer et al. (arXiv:2304.15004, 2023) respond that these emergences may be an artifact of the chosen metrics: with continuous metrics, performance scales smoothly without visible discontinuity. The question remains open.
Do LLMs understand, or simulate understanding? Bender et al. (ACM FAccT 2021) formulate the “stochastic parrot” critique: LLMs combine linguistic forms without reference to any real meaning or intention. Recent work proposes an intermediate term — contextual extrapolation from acquired knowledge — but this positioning itself remains debated.
Key takeaways
- Current LLMs are the product of five chained breakthroughs: word embeddings (2013), Transformer (2017), pre-training/fine-tuning (2018), scaling laws (2020), alignment via RLHF (2022).
- ChatGPT’s popularity is an interface breakthrough, not a model breakthrough: the LLM had existed for months; what changed public perception was non-technical access.
- The Chinchilla law (2022) corrected the raw scale race: for a fixed compute budget, it is better to train a smaller model on more data.
- RLHF showed that a model 100 times smaller can be preferred by humans over a larger model if its alignment is better.
- Whether LLMs “understand” or “simulate” remains an open question with genuine disagreement in the scientific literature.