In brief

The AIs capable of conversing with you didn’t come out of nowhere. They are the result of a decade of research, with a few pivotal moments where everything shifted. The most visible: November 2022, when ChatGPT put this technology in everyone’s hands.

Think about the history of engines: we didn’t jump from the stagecoach to the electric car in one go. There was the steam engine, then the internal combustion engine, then electronic injection — and many missteps along the way. Conversational AIs follow the same logic — each step depended on the previous one, and 2022 looks a lot like 1908 when Henry Ford released the Model T: not the invention itself, but the moment the object became accessible to the general public.


Before 2017 — Researchers teach machines to read

For a long time, computers treated words as disconnected labels. “Cat” and “dog” were as far apart as “cat” and “galaxy.” Not very useful for understanding a sentence.

In 2013, a Google team found a way to place words on a kind of map. On this map, words that are close in meaning end up close geographically. “King” and “queen” are neighbors, “car” and “truck” too. For the first time, a machine captured something that resembled meaning.

But a major problem remained: the word “bank” had the same position on the map whether you were talking about a riverbank or a financial institution. The machine didn’t take context into account. In 2018, researchers solved this problem: from then on, the same word changes its representation depending on the sentence it appears in. The machine begins to “read” for real.

In short: before 2018, every word had a fixed entry in the machine’s dictionary. After 2018, every word is read in its context — the program understands that “river bank” and “investment bank” mean two different things.


2017 — The invention that changes everything

In June 2017, a Google team published a paper with what would become a famous title: “Attention Is All You Need.” In it, they describe a new architecture — a way of organizing computations — called the Transformer.

The idea, in simple terms: when you read a sentence, you don’t read word by word in order. You go back and forth, linking “he” to the character mentioned earlier, understanding that a word at the end of the sentence sheds light on the beginning. The Transformer does the same. Instead of processing words one by one (as previous systems did), it looks at all of them simultaneously and calculates which ones are important relative to each other.

Two major consequences. First, it’s much faster — you can fully exploit the power of graphics processors (GPUs). Second, the machine finally grasps the connections between words that are far apart in a text. The Transformer is the foundation on which all current conversational AI systems are built.

In short: the Transformer is the internal combustion engine of conversational AI. Every system you’ve heard of — ChatGPT, Claude, Gemini, Mistral — runs on this same 2017 foundation. Without this piece, none of it would exist.


2018–2020 — Bigger means better (mostly)

From 2018 onward, all research labs adopted the Transformer and began training it on astronomical quantities of text. The principle: have the machine read billions of web pages, books, and articles so that it learns the structures of language. Then fine-tune it on specific tasks.

OpenAI launched GPT-1 in 2018, then GPT-2 in 2019. GPT-2 surprised everyone: without being explicitly taught, it could summarize texts, translate, and answer questions. It had simply absorbed so much text that it had “understood” how language works.

In 2020, GPT-3 pushed the needle even further: 175 billion parameters (the model’s internal settings). Give it an example or two in your prompt, and it generalizes immediately. Researchers discovered an empirical rule: the bigger the model and the more text it has read, the better it performs. The race for scale was on.

YearFlagship modelParametersObserved leap
2018GPT-1117 millionTransferable learning from a large corpus to specific tasks.
2019GPT-21.5 billionEmergent capabilities (summarization, translation) without dedicated training.
2020GPT-3175 billionFew-shot learning: one or two examples are enough to generalize.
2022GPT-3.5 (ChatGPT)~ 175 billionAlignment via human feedback: the model becomes usable day-to-day.
2023GPT-4undisclosed (trillion order)Multimodality (image), longer and more stable reasoning.

November 2022 — The ChatGPT moment

Until then, these models were impressive but frustrating. GPT-3 could write a convincing poem, then follow it with a toxic or completely fabricated response. It followed statistics, not intentions.

The key was alignment. OpenAI researchers found a way to make the model more useful: they had humans evaluate responses, then used this feedback to adjust the machine’s behavior. The result was spectacular: a model a hundred times smaller, but corrected with this method, was preferred by human testers in 85% of cases compared to raw GPT-3.

On November 30, 2022, OpenAI placed this technology in a simple chat interface, accessible to anyone. ChatGPT was born.

One million users in five days. One hundred million in two months. No tech product had ever reached this level of adoption so quickly. The technology had existed for months — what changed was that anyone could type a question and get a coherent answer, without knowing anything about computers.

This is the precise moment that shifted conversational AI from the research world to the daily lives of millions of people.


2023–2025 — The explosion and new challenges

After ChatGPT, everything accelerated. In 2023, OpenAI released GPT-4, capable of understanding images as well. Google launched Gemini, its own model. Anthropic developed Claude. Dozens of companies entered the race.

A parallel movement developed around open models, which anyone can download and use freely. Meta published LLaMA, Mistral AI (a French startup) released compact models that competed with much larger giants. The technology was democratizing.

In 2024–2025, a new stage emerged: models that “think” before responding. Instead of giving an immediate answer, they unfold a step-by-step reasoning process — much like you might when solving a complex problem on paper before giving your final answer. OpenAI, DeepSeek (China), and others published models of this type. DeepSeek showed that comparable results could be achieved with far fewer resources, challenging the assumption that only American giants could compete.

Today, conversational AI is everywhere: writing, coding, research, customer service. But fundamental questions remain open. Do these machines truly understand what they say, or do they combine words in a very sophisticated way without real meaning? Researchers disagree, and it’s a debate we probably won’t settle anytime soon.


Key takeaways

  • Conversational AI is the product of a decade of research, not a single invention.
  • The Transformer (2017) is the shared technical foundation of ChatGPT, Claude, Gemini, and all the others.
  • ChatGPT (November 2022) was not a scientific breakthrough in itself — it was the moment a powerful model became accessible to everyone through a simple interface.
  • The current race focuses on models that “reason” step by step, and on democratization through open models.
  • Whether these machines truly “understand” anything remains an open and actively debated question among researchers.