In brief
Article revised on 6 September 2026: the lineup changed shape over the summer, and the name of the most capable model is no longer the one the previous version of this page announced.
Claude is the family of language models developed by Anthropic, a laboratory founded in 2021 by former OpenAI researchers. From March 2024 to the summer of 2026, the lineup held to three tiers named in ascending order of capability: Haiku (lightweight, economical), Sonnet (cost-performance balance), Opus (the most capable). That regularity had become a way of reading the market — and it has just stopped.
Since 1 September 2026, the documented lineup has four models, and the top no longer carries the Opus name: Haiku 4.5, Claude Sonnet 5, Claude Opus 5, and Claude Fable 5.1, placed above Opus for demanding reasoning and long-horizon agentic work. A fifth, Claude Mythos 5.1, shares Fable 5.1’s specifications but is available by invitation only.
What sets Claude apart from its direct competitors comes down to two aspects: a proprietary alignment technique called Constitutional AI, and a strong positioning on long and autonomous tasks — large document analysis, agentic workflows, interface control. The model does not generate images and has no internet access without explicit tools.
In short: the three-course menu became a four-course one. Haiku is still the starter — fast, cheap. Sonnet is still the main course — balanced, good value. Opus is no longer the chef’s special but the default choice: it is the one Anthropic itself recommends starting from. Above it appeared a tasting dish, Fable, twice the price and slower, reserved for cases where Opus pushed hard is not enough. The practical lesson is not the name: it is that a naming convention stable for two and a half years can vanish in one summer, and that code naming a model in hard-coded form ages faster than you think.
Identity card
| Field | Value |
|---|---|
| Organization | Anthropic (San Francisco, founded 2021) |
| First version | Claude 1, March 2023 |
| Type | Multimodal language model (text + vision) |
| Access | API (platform.claude.com), claude.ai interface, Enterprise plans, Amazon Bedrock, Google Cloud, Microsoft Foundry |
| Context window | 1 million tokens on Fable 5.1, Opus 5 and Sonnet 5 — 200,000 on Haiku 4.5 |
| Lineup as of 6 September 2026 | Haiku 4.5 · Sonnet 5 · Opus 5 · Fable 5.1 (+ Mythos 5.1 by invitation) |
History
Claude 1 (March 2023). Public launch in two variants: Claude (flagship) and Claude Instant (faster, less expensive). Stated positioning: “helpful, honest, and harmless”. Available via API and claude.ai.
Claude 2 (July 2023). Improvements in coding, mathematics, and reasoning. Version 2.1 (November 2023) introduced the 200,000-token context window, targeting enterprise use cases — legal, finance, document research.
Claude 3 (March 2024). First use of the Opus / Sonnet / Haiku naming scheme. Introduction of vision capabilities: image analysis, charts, PDF documents. Three clearly distinct performance and pricing tiers.
Claude 3.5 Sonnet (June 2024). A mid-tier model that outperforms the flagship from the previous generation. Shifts developer perception around model selection. The v2 release (October 2024) introduced Computer Use — the first AI model to directly control a computer interface.
Claude 3.7 Sonnet (February 2025). Introduction of extended thinking: hybrid reasoning that allows the model to “think step by step” before responding, with a configurable thinking token budget.
Claude 4 — Opus 4 and Sonnet 4 (May 2025). SWE-bench Verified: Opus 4 at 72.5% (79.4% in high-compute mode), Sonnet 4 at 72.7%. Hybrid models combining instant response and extended thinking. Opus 4 capable of operating continuously for several hours. Opus 4 pricing: $15/1M input tokens, $75 output.
Claude 4.5 (2025). Sonnet 4.5 (September 2025): SWE-bench 77.2%, OSWorld 61.4%, GPQA Diamond 83.4%, AIME Python 100%. Autonomous operation documented over more than 30 hours. Haiku 4.5 (October 2025): performance approaching Sonnet 4 at one-third of the cost ($1/$5 per million tokens). Opus 4.5 (November 2025): improvements in coding and office tasks, problematic response rate of 0.22% on single-turn violative queries.
Claude 4.6 (February 2026). Sonnet 4.6 (February 17): SWE-bench 79.6%, OSWorld 72.5%, first Sonnet model preferred over the Opus of the previous generation in coding evaluations (70% of cases vs. Sonnet 4.5, 59% vs. Opus 4.5). Introduces the 1M token context window in beta. Opus 4.6 (February 5): addition of agent teams, PowerPoint integration. Pricing for both: $5/$25 — a 66% reduction compared to Opus 4.
Opus 4.7 and 4.8 (2026). Two intermediate iterations, now classed as legacy but still served. Opus 4.7 brings a quiet change with loud consequences: a new tokenizer. For the same token count, the window now holds about 555,000 words instead of 750,000 — the counter did not change size, the unit changed value. A context budget calibrated on the old tokenizer has to be recomputed.
Generation 5 (July–September 2026). Sonnet 5 and Opus 5 replace the 4.x line. Claude Opus 5 shipped on 24 July 2026 (claude-opus-5, 1M context, 128K output, $5/$25), with two breaking changes for code written for Opus 4.8: thinking is on by default, and can only be disabled at effort high or below.
Claude Fable 5.1 (1 September 2026). A new top of the range, above Opus (claude-fable-5-1, 1M context, $10/$50). Anthropic explicitly recommends starting with Opus 5 and moving to Fable 5.1 only for demanding reasoning and long-horizon agentic work, or when evaluations on Opus 5 at higher effort still fall short. Its cache reads are billed at $0.25 per million tokens — 2.5% of the input price, against 10% across the rest of the lineup, which changes the economics of long conversations far more than the headline price suggests. Claude Mythos 5.1 offers the same capabilities, by invitation only (Project Glasswing).
In short: the trajectory comes down to three movements. SWE-bench Verified — real-bug resolution — went from 64% in June 2024 to 79.6% in February 2026, a 15-point gain in twenty months; beyond that, Anthropic no longer publishes this comparison on its model pages, and this page does not invent it. The price of the reference model fell from $75 to $25 per million output tokens between Opus 4 and Opus 5, that is 66% less for a more capable model — but a new upper tier appeared at $50, so the top of the range actually got dearer. And the context went from 100,000 tokens to 1 million as standard, no longer in beta: roughly 555,000 words held at every request.
Capabilities
The lineup as of 6 September 2026 (vendor documentation)
| Model | Context | Max output | Input / output price | Thinking | Knowledge cutoff | Retirement no sooner than |
|---|---|---|---|---|---|---|
| Claude Fable 5.1 | 1M | 128K | $10 / $50 | adaptive, always on | Jun 2026 | 1 Sep 2027 |
| Claude Opus 5 | 1M | 128K | $5 / $25 | adaptive | May 2026 | 24 Jul 2027 |
| Claude Sonnet 5 | 1M | 128K | $2 / $10 | adaptive | Jan 2026 | 30 Jun 2027 |
| Claude Haiku 4.5 | 200K | 64K | $1 / $5 | extended | Feb 2025 | 15 Oct 2026 |
Prices in dollars per million tokens. Cache reads cost 10% of the input price, except on Fable 5.1 where they drop to 2.5% ($0.25). Batch requests are discounted 50%.
Two columns deserve a professional reader’s attention, because they were not on vendor sheets two years ago.
The knowledge cutoff, first: it is now published per model, and it spans sixteen months between the newest and the oldest of the lineup. Haiku 4.5 stops in February 2025 while Fable 5.1 runs to June 2026. Two models of the same brand, served on the same day, do not share the same idea of what the present is.
Retirement no sooner than, second: Anthropic commits to a date before which a model will not be withdrawn. That is governance information, not performance — it says how long an integration can live untouched. Haiku 4.5 is the most exposed: its date falls in October 2026.
Legacy models still served: Fable 5, Opus 4.8, 4.7, 4.6, 4.5, Sonnet 4.6 and 4.5.
Thinking changed mode
Up to generation 4.6, thinking was extended: the developer switched on a mode and set a token budget. Since then it is adaptive: the model decides its own depth, steered by an effort parameter defaulting to high. The manual mode is deprecated on Opus 4.6 and Sonnet 4.6, and refused on later models. On Fable 5.1, adaptive is the only mode available.
This is not an API detail: the dial for “how much the model thinks before answering” moved from the developer’s hand to the model’s, the developer keeping only a statement of intent. Code that hard-set a thinking budget no longer works as written.
Published benchmarks (vendor data, generations 3.5 to 4.6)
| Model | SWE-bench | OSWorld | GPQA Diamond |
|---|---|---|---|
| Claude 3.5 Sonnet | 64% (internal) | — | — |
| Claude Opus 4 | 72.5% (79.4% high-compute) | — | — |
| Claude Sonnet 4 | 72.7% | — | — |
| Claude Sonnet 4.5 | 77.2% | 61.4% | 83.4% |
| Claude Sonnet 4.6 | 79.6% | 72.5% | — |
SWE-bench Verified measures the resolution of real bugs in open-source repositories. OSWorld evaluates interface control on an operating system under real conditions.
Documented strengths
Coding and bug resolution. Claude has held top positions on SWE-bench across several generations. Sonnet 4.6 reaches 79.6%; Sonnet 4.5 had scored 100% on AIME Python.
Multi-step reasoning. Extended thinking (since 3.7) enables a configurable reasoning budget before the final response. Useful for mathematical, logical, or engineering problems.
Long document analysis. The 1-million-token context moved from beta to standard on Fable 5.1, Opus 5 and Sonnet 5 — roughly 555,000 words, an entire documentation corpus or a complete code base. Mind the recomputation: the tokenizer introduced with Opus 4.7 fits fewer words into the same token count than before (555,000 against roughly 750,000).
Interface control. Computer Use (since 3.5 Sonnet v2) allows Claude to control a browser or desktop. OSWorld 72.5% on Sonnet 4.6 makes it the highest-performing model measured on this task at that date.
Vision. Available since Claude 3. Up to 600 images or PDF pages per request. Interpretation of charts, graphs, and tables.
In short: three use cases stand out where Claude clearly distinguishes itself. First, production code — 79.6% of SWE-bench Verified instances are resolved under the benchmark’s published conditions. That is not 79.6% of real bugs: the test set is drawn from tickets within a chosen scope, with an automatable success criterion, which matches neither the difficulty nor the shape of a bug met in production. Second, large-scale document analysis — 200k tokens cover an entire 500-page specification, and the 1M beta absorbs a full codebase. Third, interface control — Sonnet 4.6 reaches 72.5% on OSWorld, one of the few benchmarks measuring an agent against a real operating system rather than a web page.
Choosing the right Claude tier
| Context | Recommendation | Why |
|---|---|---|
| General case, including coding and enterprise work | Opus 5 | The starting point Anthropic itself recommends. $5/$25, 1M context, adaptive thinking. |
| Lightweight conversational task, high-volume support, hard budget constraint | Haiku 4.5 | $1/$5, the lowest latency in the lineup. 200K context and knowledge stopping in February 2025: rule it out as soon as recency matters. |
| Standard analysis, writing, medium-complexity code | Sonnet 5 | $2/$10 for the same 1M context as the top of the range. The best volume-per-euro in the lineup. |
| Demanding reasoning, long-horizon agentic work, or measured failure of Opus 5 at higher effort | Fable 5.1 | $10/$50 and slower — the premium is justified only after measuring that Opus 5 falls short, not before. Cache reads at $0.25 make long conversations cheaper than they look. |
| Very long conversation with heavily reused context | Fable 5.1, worth computing | Cache reads cost 2.5% of the input price there against 10% elsewhere: over a context re-read dozens of times, the headline price gap sometimes reverses. |
| Image generation output or web search without tools | Out of scope for Claude | Capabilities natively absent — point users to OpenAI or Gemini depending on need. |
Known limitations
No image generation. Claude analyzes images as input but does not produce images as output. This capability is natively absent, with no announced roadmap.
Hallucinations. The rate varies by task: approximately 4.4% on standard document summarization, around 10% on difficult benchmarks for Opus 4 (source: Suprmind, 2026). Models in extended thinking mode can “over-think” and drift from the source material. Constitutional AI encourages the model to signal its uncertainty rather than hallucinate with confidence — which reduces confident errors, but not total errors.
Knowledge cutoff. It is now published model by model, which was not the case until 2026: June 2026 for Fable 5.1, May 2026 for Opus 5, January 2026 for Sonnet 5, February 2025 for Haiku 4.5. The sixteen-month spread inside a single lineup is itself a usage limit: the cheapest model is also the one that knows the least. Without a web search tool, Claude has no information later than that date.
Errors on long code. Errors have been documented when generating a single very long code block. The recommended practice is to break tasks into steps.
No native internet access. Except via external tools (MCP servers, web search integrated depending on plan). Claude operates on its frozen knowledge by default.
Key takeaways
The three-tier naming convention, stable from March 2024 to the summer of 2026, has stopped being stable. The lineup has four models and the most capable one is no longer called Opus. For anyone writing code or an architecture note, the lesson is more useful than the name itself: a hard-coded model identifier is a disguised expiry date.
The progression remains measurable on the published part: SWE-bench goes from 64% (3.5 Sonnet) to 79.6% (Sonnet 4.6). Beyond that, Anthropic no longer publishes this comparison on its model pages — so this page stops there rather than quoting aggregators.
The objective strengths remain coding, long document analysis, and interface control. The structural limitations remain the absence of image generation and a non-negligible hallucination rate on difficult tasks. The knowledge cutoff, however, is no longer an unknown: it is published, and its sixteen-month spread between the top and bottom of the lineup has become a selection criterion in its own right.
The price of the reference model fell 66% between Opus 4 ($15/$75) and Opus 5 ($5/$25). But a higher tier appeared above it at $10/$50: the market is not falling, it is spreading out. And what really changed the economics of use is not the headline price but the cache read, down to 2.5% of the input price on Fable 5.1 — on a long conversation re-read dozens of times, that is the line item that dominates the bill.