In brief

The GPT family is the line of language models developed by OpenAI since 2018. It laid the foundations of the current large language model market: GPT-3 showed that a sufficiently large model could generalize across a wide variety of tasks, ChatGPT brought these capabilities to the general public, and the o series introduced a new way of thinking about inference — no longer as simple token-by-token generation, but as an extended reasoning process.

Two branches long coexisted: GPT models (4o, 4.1) optimized for versatility and speed, and o models (o1, o3, o4-mini) investing additional compute at inference time. They converged in GPT-5 (August 2025), whose selling point was an automatic router: the user no longer had to choose.

As of 6 September 2026, the choice is back — under other names. The documented catalogue puts GPT-6 Astra at the top, with three named GPT-5.6 variants below it: Sol, Terra, Luna, separated by a factor of sixteen on output price — and more than forty between the cheapest and the top of the range. Alongside them sit specialised lines the previous generation did not have: image, realtime and voice, transcription, and a family dedicated to cybersecurity.

In short: imagine two assistants at a bank counter. The first (GPT-4o, 4.1) answers off the cuff — fast, fluent, ideal for routine questions. The second (o1, o3) sets your file aside, reads through it, takes notes, double-checks before coming back with a substantiated answer — slower but more reliable on complex cases. GPT-5 places the two assistants side by side with a dispatcher that routes the request to the right profile based on detected difficulty. The whole thing is a unified product for the end user.


Identity card

FieldValue
OrganizationOpenAI
First versionGPT-1 (2018)
TypeDecoder-only Transformer, pre-trained and instruction-tuned
AccessAPI (OpenAI), ChatGPT interface (free and Pro subscription)
Context window1,050,000 tokens on GPT-6 Astra and the GPT-5.6 lineup — 128,000 (GPT-4o, historical)
Lineup as of 6 September 2026GPT-6 Astra · GPT-5.6 Sol · GPT-5.6 Terra · GPT-5.6 Luna (+ image, realtime, transcription, cyber lines)

History

2018–2022: the foundations

GPT-1 (2018) introduces the decoder-only Transformer architecture, pre-trained on a corpus of books. GPT-2 (2019) scales this approach further, with an initial staged release by OpenAI due to concerns over misuse. GPT-3 (2020), with its 175 billion parameters, reveals the emergence of unexpected generalization capabilities: the model solves tasks it was never explicitly trained for. GPT-3.5 / ChatGPT (November 2022) adds instruction fine-tuning (RLHF) and brings broad public access.

2023: GPT-4

Launched on March 14, 2023, GPT-4 significantly improves reliability and problem-solving. Its internal architecture has never been officially confirmed, though third-party sources suggest a Mixture of Experts organization. OpenAI stopped publishing parameter counts from this version onward.

2024: multimodality and reasoning

GPT-4o (May 2024) is the first natively multimodal GPT model: it processes text, audio, and images in a unified pipeline, without an intermediate transcription step. Audio response latency reaches 320 ms, comparable to human reaction time. Context window: 128,000 tokens.

o1-preview (September 2024) inaugurates the o series. These models are trained to “think” before responding: an internal chain-of-thought process, invisible to the user, explores multiple paths before formulating an answer. This is not simple CoT prompting — the reasoning is integrated into training and inference (“test-time compute”).

In short: the o1 break is not a new bigger model — it is a new way of using compute at inference. Where GPT-4o produces an answer in 2-3 seconds on average, o1 can take 30 seconds to several minutes to mentally explore multiple paths before formulating the final answer. This “thinking” multiplies compute cost by 5× to 50× depending on difficulty. The bet: on mathematical and scientific problems, this overhead is justified by a massive quality leap (74% AIME vs 12% for GPT-4o, that is ×6).

2025: extended contexts and unification

GPT-4.1 (April 2025) extends the context window to 1 million tokens. o3 and o4-mini (April 2025) extend reasoning capabilities to tools (web search, Python execution, image generation), with a 200,000-token window and 100,000 tokens of output.

GPT-5 (August 2025) unifies the two branches: an automatic router distributes requests between a fast mode and a thinking mode based on problem complexity.

In short: the GPT-5 convergence is also an economic convergence. Before, the user had to explicitly choose between GPT-4o ($1-2/M output tokens) and o3 ($8/M output tokens). Having to pick the right model is a cognitive load — and a risk of overpaying for a simple question. GPT-5 abstracts this choice by routing requests automatically: 80-90% of prompts do not require thinking mode and stay economical, the 10-20% complex ones switch to thinking with no user intervention.

2026: retiring the old, then bringing choice back

On February 13, 2026, OpenAI removed GPT-4o, GPT-4.1, GPT-4.1 mini, and o4-mini from ChatGPT. These models remain available via the API.

The GPT-5.6 lineup replaces the single model with three named variants, all at a 1.05-million-token context: Sol for complex professional work ($4/$20 per million tokens, aliased gpt-5.6), Terra balancing capability and expense ($2/$12), Luna for cost-sensitive workloads ($0.20/$1.20). Between Luna and Sol, output price varies by a factor of sixteen; between Luna and the top of the range, by more than forty.

GPT-6 Astra (3 September 2026) takes the head of the catalogue — “our most capable model, built for the hardest end-to-end work,” model ID gpt-6-astra, 1.05-million-token context, 128,000-token output, $10/$50. The rollout is phased: limited preview for trusted partners on 3 September, opening to paid tiers the next day in a restricted version that refuses certain prompts, notably in cybersecurity. Companies enrolled in OpenAI’s cyber programme were served first.

A cybersecurity line appears in the catalogue — GPT-5.6 Cyber, Daybreak Red, Daybreak Blue — where earlier generations exposed no model dedicated to the domain. It is the mark of a market segmenting by profession rather than by size alone.

In short: GPT-5’s selling point in 2025 was “you no longer choose your model, the router does it.” A year later, the catalogue documents four named general models and half a dozen specialised ones — choice is back, moved from a technical setting to a commercial name. For anyone deciding an architecture, the practical question has not changed: between Luna and Astra, output price is multiplied by more than forty, and nothing in the name says where the gap starts being worth it. That is measured, not read off a datasheet.


Capabilities

Benchmarks

o1 series (September 2024)

Benchmarko1GPT-4o (reference)
AIME 202474% (1 attempt) / 93% (1000 attempts, reranking)12%
GPQA Diamond78.1%—

The GPQA Diamond score of 78.1% exceeds the median level of human doctoral experts in their field. Claude 3.5 Sonnet reached 67.2% on the same benchmark.

Source: OpenAI — Learning to reason with LLMs

o3 and o4-mini (April 2025)

Benchmarko3o4-mini
AIME 2025 (with Python)98.4%99.5%
SWE-bench Verified69.1%68.1%
GPQA Diamond87.7%81.4%

Source: OpenAI — Introducing o3 and o4-mini

GPT-5 (August 2025)

BenchmarkScore
AIME 2025 (no tools)94.6%
SWE-bench Verified74.9%
Aider Polyglot88%
MMMU84.2%

Source: OpenAI — Introducing GPT-5

Multimodality

GPT-4o accepts text, audio, images, and video (converted to frames at 2–4 fps, without an audio track) as input. Output: text, audio, image. Audio processing is native — no separate transcription pipeline — which reduces latency and improves understanding of prosodic nuance.

The o3 and o4-mini models add deep visual reasoning and integration of ChatGPT tools into the chain of thought.


Known limitations

Hallucinations

GPT-4o inherits the hallucination problems of previous generations. High-fidelity voice generation introduces an additional risk: the fluency of the output can induce poorly calibrated confidence in users. OpenAI’s audio transcription model produces approximately 90% fewer hallucinations than Whisper v2 in internal tests in noisy environments — but this figure concerns transcription, not content generation.

Source: GPT-4o System Card

Knowledge cutoff

ModelCutoff
GPT-4o (initial version)October 2023
GPT-4o (updated version)June 2024
GPT-4.1, 4.1 mini, 4.1 nanoJune 2024

The cutoff for o3 and o4-mini has not been confirmed in the primary sources consulted, and OpenAI’s models page publishes none for the 5.6 lineup or for Astra either. That is a notable difference in practice from Anthropic, which now publishes a knowledge cutoff per model: for comparable catalogues, the information is not available on both sides.

API costs

Reasoning models are significantly slower than standard models: the internal “thinking” time increases latency considerably. Pricing reflects this investment:

Catalogue as of 6 September 2026 (OpenAI documentation):

ModelContextInput ($/M tokens)Output ($/M tokens)Stated purpose
GPT-6 Astra1.05M$10.00$50.00the hardest end-to-end work
GPT-5.6 Sol (gpt-5.6)1.05M$4.00$20.00complex professional work
GPT-5.6 Terra1.05M$2.00$12.00capability / expense balance
GPT-5.6 Luna1.05M$0.20$1.20cost-sensitive workloads

Previous generation, still served by the API:

ModelInput ($/M tokens)Output ($/M tokens)
GPT-4.1$2.00$8.00
GPT-4.1 mini$0.40$1.60
GPT-4.1 nano$0.10$0.40
o3$2.00$8.00
o4-mini$1.10$4.40

Discounts apply via the Batch API (–50%, results within 24h) and prompt caching (cached tokens at –50%).

The striking fact is not the top price but the spread: at an identical 1.05-million-token context, output runs from $1.20 to $50 per million tokens depending on the variant. The same volume of context therefore costs forty times more to process from one end of the catalogue to the other — capability is what is billed, not window size.

Discounts apply through the Batch API (–50%) and prompt caching.

Source: OpenAI — Models

GPT-6 Astra announcement date: CNBC, “OpenAI announces rollout of GPT-6 Astra model”, 3 September 2026 (unlinked reference: the site refuses automated requests, so a link would be reported dead by our checker).

Architectural opacity

Since GPT-4, OpenAI no longer publishes parameter counts or architectural details of its models. The Mixture of Experts architecture often cited for GPT-4 has never been officially confirmed. This opacity makes rigorous comparison with open-weight models difficult.

Which GPT model for which use

ContextRecommendationWhy
Consumer app, latency-critical, tight budgetGPT-5.6 Luna$0.20/$1.20 for the same 1.05M context as the top of the range. The cheapest option in the current catalogue.
Standard office task, product dialogue, routine writingGPT-5.6 Terra$2/$12 — the stated balance point between capability and expense.
Complex professional work, long analysis, codeGPT-5.6 Sol$4/$20; the gpt-5.6 alias points here, so it is the implicit default if you did not choose.
Hard end-to-end task, long-horizon agentic workGPT-6 Astra$10/$50, phased rollout and a restricted version on certain topics. Reserve it for cases where the 5.6 variants have been measured insufficient.
Math, scientific, or complex code problem requiring accuracyo3 or o4-mini98.4% AIME 2025, 87.7% GPQA. The extra cost ($8/M output tokens) is justified by ×6 quality jump vs GPT-4o.
Very long document analysis (report, codebase)GPT-5.6 LunaA 1.05M-token window — roughly a 700-page book — at $0.20/$1.20. The entire current catalogue holds that window; GPT-4.1 (1M, $2/$8) is still served but costs ten times more on input and nearly seven times more on output, for a shorter context.
Native multimodal (real-time audio, video)GPT-4oAudio pipeline without intermediate transcription, 320 ms latency — comparable to human reaction time.

Key takeaways

  • The GPT family has undergone three major shifts: quantitative (scaling), qualitative (instruction tuning), and algorithmic (inference-time reasoning with the o series).
  • GPT-5 represents the convergence of these directions: an automatic router distributes requests between fast mode and thinking mode based on demand.
  • Reasoning models (o1, o3) invest additional compute at inference time — this is not prompting, it is integrated into training.
  • The exact architecture, real parameter count, and training data have remained opaque since GPT-4: benchmarks are the main external observation window.