In brief
Choosing a language model for professional use takes four facts: the size of the context window, the price, the date at which the model’s knowledge stops, and how long it will keep being served. The first three decide what the model can do and what it costs. The fourth decides the lifespan of whatever you build around it.
On 6 September 2026, those four facts were looked up on the same day, on the official model pages of Anthropic, OpenAI and Google. None of the three publishes all four. The gap between the most complete and the most reticent is not a technical-writing detail: it changes the nature of the scoping work, and it determines what an architect can assert independently rather than on trust.
In short: picture three car dealers. The first shows, on one sheet, engine size, price, model year, and a written commitment that parts will be stocked until a given date. The second shows engine size and price, but not the year. The third gives you the car names and refers you to a salesperson for the rest. All three sell good cars. Only one lets you work out, on your own, how long yours will keep running.
The method, so the survey can be redone
There is no methodological mystery here, and that is deliberate: the value of this survey lies in anyone being able to redo it and find the same thing — or find that it has changed.
Three pages, consulted on 6 September 2026:
- the “Models overview” page of Anthropic’s documentation;
- the “Models” page of OpenAI’s developer documentation;
- the “Models” page of the Gemini API documentation.
For each vendor, a single question: does this page — the one the documentation points to first — publish the field? Not “does the information exist somewhere on the site” — it usually does, in a release note, a blog post, or a per-model sheet. The question is the one a hurried reader comparing options actually faces: is the field where comparison is meant to happen?
The survey
| Information published on the models page | Anthropic | OpenAI | |
|---|---|---|---|
| List of current models | yes | yes | yes |
| Context window | yes | yes | no |
| Max output tokens | yes | yes | no |
| Input / output price | yes | yes | no |
| Cache read price | yes | partial | no |
| Knowledge cutoff date | yes | no | no |
| Earliest retirement date | yes | no | no |
| Model IDs on third-party platforms | yes | no | yes |
| Legacy models still served | yes | yes | yes |
Survey of 6 September 2026. A “no” means the field is absent from the comparison page, not that it cannot be found elsewhere.
What each field lets you decide
The context window is the most-watched field, and the least treacherous — with one caveat we return to below. It says how much material fits in a request.
The price is deceptively simple. The headline input and output prices say nothing about the line item that actually dominates a conversational agent’s bill: the cache read, that is, what it costs to re-read the same context at every turn. At Anthropic it runs at 10% of the input price across the lineup, and 2.5% on the top model. Over a conversation re-read fifty times, that gap weighs more than the headline price — enough, sometimes, to reverse the ranking between two models.
The knowledge cutoff says what the model does not know. It is a capability limit disguised as a footnote. At Anthropic it spreads over sixteen months inside a single lineup: June 2026 for the top model, February 2025 for the cheapest. Two models from the same vendor, served on the same day, do not share the same idea of what the present is. A system that switches automatically from one to the other to save money also switches, silently, from one year of knowledge to another.
The earliest retirement date is the rarest field and the most structuring. It says nothing about performance: it says how long the integration you write today will run before it has to be reworked. It is the only one of the four that speaks about your calendar rather than the model’s. One of the three vendors publishes it.
In short: the first three fields answer “will it work?”. The fourth answers “for how long?”. That is the question people forget to ask at decision time, and the only one that cannot be answered after the fact.
Three traps a datasheet does not flag
The name no longer orders capability
The market’s implicit convention — a top tier, a middle tier, an economy tier, ordered by name — stopped being reliable in 2026, and at two vendors at once.
At Anthropic, the most capable model is no longer from the Opus line: a new name moved above it, and the documentation explicitly recommends starting with Opus and moving up only when measurement justifies it. At Google, the model presented as the most advanced is a Flash, in stable release, while the Pro tier sits at a numerically lower version, in preview. A reader applying “Pro > Flash” picks the wrong model.
At OpenAI the movement is the reverse, and just as disorienting: the single auto-routing model of 2025, sold on the promise that you would no longer have to choose, has been replaced by three named variants between which you must choose again — separated by a factor of sixteen on output price, at an identical context window.
The phenomenon runs beyond the three vendors surveyed here. At Mistral, the model documented as the most capable is called Medium, not Large — and it is the only one of the three main models that is commercial, while those named Large and Small are open under Apache 2.0. “The most open” and “the most capable” no longer name the same model there, which neither name hints at.
The same token count no longer holds the same text
A tokenizer change shows up on no product sheet, because no field carries it. It nonetheless changes the value of every other field.
At Anthropic, the tokenizer introduced during 2026 fits roughly 555,000 words into a million tokens, where earlier models fitted roughly 750,000. The advertised window did not move; its real capacity fell by nearly a quarter. A context budget calibrated before the change, and never recomputed, overflows in silence.
A missing field is not missing data
A vendor that does not publish context windows on its comparison page does not have models without context windows. The information exists — in a per-model sheet, a release note, a blog post. What changes is the cost of verification: where a table allows a decision in five minutes, its absence imposes a page-by-page collection, to be redone at every lineup revision, with no guarantee that it is current.
That is why the difference is real rather than rhetorical. The most forthcoming vendor is not necessarily the most capable. It is the most checkable — and for an architecture decision you will have to defend in six months, checkability is a property of the vendor in the same way latency is.
What this changes concretely for an architecture decision
Pin the identifier, never the tier. Writing “the vendor’s top-end model” in a specification is writing a moving target. The two 2026 examples make the point: a hard-coded identifier is a disguised expiry date, but a tier written in plain words is worse — it changes meaning without notice.
Date your survey, and feed the date back into the documentation. A model comparison is only true on a date. This article carries its own in its section headings and its sources; an internal architecture note should do the same. Undated, a comparison table becomes a source of error within a quarter.
Measure what is not published rather than estimating it. When the knowledge cutoff is not given, a handful of questions about dated facts bounds it better than a guess. When the real capacity of a window is uncertain, a text of known length measures it in one request. These checks take minutes and replace a belief with a number.
Ask in writing for what is missing. A retirement date is contractual information before it is technical. That only one vendor publishes it unprompted does not mean the others refuse to commit — it means you have to ask, and that the answer belongs in the decision file.
What to do when a field is missing
| Context | Recommendation | Why |
|---|---|---|
| The retirement date is not published | Ask for it in writing before committing, and file the answer | It is the only fact that bounds the lifespan of your integration. It is contractual before it is technical: not publishing it is not a refusal to commit. |
| The knowledge cutoff is not published | Bound it by measurement: a handful of questions about dated facts | A few minutes replace a guess with an interval. Redo it at every model change, including a tier change within the same vendor. |
| The context window is not on the comparison page | Go and find it per model, and date the survey | The information exists elsewhere. What costs is re-collecting it at every lineup revision — so write the date next to the number. |
| The tokenizer changed since you calibrated | Re-measure real capacity with a text of known length | An unchanged advertised window may have lost nearly a quarter of its capacity. No datasheet field flags it. |
| You are writing a specification or an architecture note | Pin the exact identifier, never the tier name | Tiers stopped ordering capability in 2026 at three vendors at least. An identifier ages; a tier changes meaning without notice, which is worse. |
| Conversational use with context re-read every turn | Compare on the cache read price, not the headline price | Its rate varies by a factor of four inside one catalogue, and it is what dominates the bill. |
Key takeaways
- Four facts decide an architecture: context window, price, knowledge cutoff, service lifespan. None of the three major vendors publishes all four on its comparison page, as of 6 September 2026.
- The rarest is also the most structuring: the earliest retirement date is the only one that speaks about your calendar rather than the model’s. One vendor gives it.
- The knowledge cutoff varies by sixteen months inside a single lineup. Switching to a cheaper model also means switching to one that knows less — and nothing in the price says so.
- The headline price is not the price paid: in conversational use, the cache read dominates, and its rate varies by a factor of four inside a single catalogue.
- Tier names stopped ordering capability in 2026, at three vendors at least — and at one of them the most capable model is also the only one that is not open. Pin a precise identifier, and date it.
- A tokenizer change can cut the real capacity of an unchanged advertised window by a quarter. No datasheet field flags it.