In brief

Cohere is a Canadian company founded in 2019 in Toronto that develops large language models exclusively for enterprises. Unlike OpenAI or Anthropic, it offers no consumer-facing equivalent of ChatGPT: its sole customer is the B2B market, and its models are designed to integrate into organizations’ private infrastructures — VPC, on-premises, sovereign cloud.

One of its co-founders, Aidan Gomez, contributed as a teenager to the foundational paper on the Transformer architecture. This technical origin is reflected in the product strategy: Cohere bets on native RAG (the model automatically cites its sources), multilingualism (up to 101 languages through its research lab), and computational efficiency (models requiring 2 GPUs where competitors demand 8).

As of September 2025, Cohere is valued at approximately 7 billion dollars with annualized revenues of 150 million dollars, and is preparing an initial public offering.


Identity card

CriterionValue
Founded2019, Toronto (Canada)
HeadquartersToronto, Ontario, Canada
Staff~843 people (2025)
StatusPrivate company — IPO announced (horizon 2026)
FoundersAidan Gomez (CEO), Nick Frosst, Ivan Zhang
Valuation~7 Bn USD (September 2025)
Revenue240 M$ annualized (February 2026)
Total raised~1.6 Bn$ (2019–2025)

History — from Google Brain to Toronto

2019 — A founding rooted in Transformer research

Cohere was born in 2019 from a straightforward idea: the enterprise market needs sovereign LLMs, with no competing product, deployable within its own infrastructure. The three founders share a common background — the University of Toronto and the Google Brain ecosystem.

Aidan Gomez (CEO) is one of the eight co-authors of “Attention Is All You Need” (Vaswani et al., NeurIPS 2017), the foundational paper on the Transformer architecture. He was 20 years old at the time, working as an intern at Google Brain. Nick Frosst, a researcher at Google Brain, and Ivan Zhang, also from that ecosystem, complete the founding trio.

Toronto, not San Francisco — this choice is deliberate. The Canadian ecosystem benefits from the University of Toronto (where Geoffrey Hinton trained a generation of researchers), immigration-friendly policies for AI talent, and an active government in supporting national digital sovereignty.

In plain terms: Cohere was founded by researchers from the group that invented the architecture underlying virtually all current LLMs. The choice of Canada reflects a deliberate strategy of geographic differentiation and sovereign positioning.

2021-2024 — Growth by stages

Series A (40 M$, 2021) attracted names validating its technical credibility: Geoffrey Hinton, Fei-Fei Li, Pieter Abbeel, and Raquel Urtasun are among the investors. Series B (125 M$, 2022, Tiger Global lead) followed quickly.

In June 2022, Cohere created Cohere For AI — a distinct, non-profit research entity, led by Sara Hooker (ex-Google Brain). This separation is structural: it allows the commercial company to focus on paid enterprise models, while the research lab produces open-weight models on multilingualism (the Aya project). In April 2025, Cohere For AI was renamed Cohere Labs.

Series C (270 M$, June 2023, valuation 2.2 Bn$) came with a strong signal: Oracle, Salesforce, and Nvidia are co-investors. The Oracle partnership is the deepest — Oracle became an investor, priority distributor on Oracle Cloud Infrastructure (OCI), and an experimentation ground for enterprise models.

2024 — The Command R year and Series D

March 2024 marked the launch of Command R (35B parameters), the first model in the R series optimized for RAG and tool use, with 128,000 token context. A month later, Command R+ (104B) positioned Cohere in the high-performance segment.

In July 2024, Series D reached 500 M$ (valuation confirmed at 6.8 Bn$), with a remarkable list of industrial investors: AMD, Nvidia, Cisco, Fujitsu, Salesforce Ventures, Export Development Canada, PSP Investments (Canadian public sector pension fund). Simultaneously, a strategic partnership with Fujitsu was announced to co-develop Takane, a Japanese enterprise LLM.

In December 2024, the Canadian government announced a 240 M$ investment — a signal of the geopolitical dimension of AI sovereignty.

2025 — Agentic and IPO preparation

March 2025: Command A (111B, 256K context) was launched as an “agentic enterprise” model, designed for autonomous agents and multi-step workflows. August 2025: Command A Reasoning, Cohere’s first reasoning model, with a configurable token budget to balance depth versus speed.

May 2026: Command A+, released on 20 May. It is the company’s first Mixture of Experts model, multimodal and multilingual, agent-oriented — and Cohere claims it runs on two H100 GPUs. The claim is worth noting: it is not a benchmark result but a deployment constraint, which is precisely Cohere’s sales argument since 2024. A model a bank can host itself is worth more, to that bank, than a model two points higher in a leaderboard but impossible to install behind its firewall.

Sara Hooker left Cohere Labs in September 2025 after three years leading the research lab. She was replaced by Marzieh Fadaee. In October 2025, Aidan Gomez publicly announced IPO preparations, with annualized revenues of 150 M$ at that date.

By February 2026, annualized revenues reached 240 M$.


Products — a range structured around enterprise use cases

Avoiding confusion: Cohere has no connection to Coherent (optical components and semiconductors manufacturer) or Cohesity (backup and data management solutions). These are entirely separate companies in different sectors.

Command R series — RAG-optimized

The Command R series is Cohere’s main product line. All these models share a RAG-native design: they produce responses with inline citations that explicitly reference the source chunks provided in context. This functionality is integrated at the model level, not added through post-processing.

Command R (March 2024, 35B parameters)

  • 128,000 token context, 10 enterprise languages (EN, FR, ES, IT, DE, PT, JA, KO, AR, ZH)
  • Optimized for RAG pipelines and tool use with low latency, high throughput
  • Available via Cohere API, AWS Bedrock, Azure AI Studio, Google Cloud Vertex AI

Command R+ (April 2024, 104B parameters)

  • Flagship high-performance enterprise model
  • Advanced RAG, tool use, multi-step reasoning
  • According to Cohere’s internal evaluations, outperforms GPT-4 Turbo and Claude 3 on RAG and tool use benchmarks [self-published, not independently replicated — no independent replication found in sources reviewed]
  • August 2024 update: +50% throughput, −20% latency, GPU footprint halved

Command R7B (December 2024, 7B parameters)

  • Most compact in the R series, optimized for high-frequency deployments (chatbots, code assistants)
  • 128,000 token context, 23 languages
  • Open weights available on Hugging Face
  • Cohere benchmarks (self-declared): first on IFeval, BBH, GPQA, MuSR, MMLU in its size class [not independently replicated]

In plain terms: the Command R series is designed to connect an LLM to a company’s document base. The model doesn’t respond from memory — it reads the provided documents and indicates where each piece of information comes from. This source transparency is a concrete business value for regulated sectors.

Command A series — agentic enterprise

Command A (March 2025, 111B parameters)

  • Designed for enterprise autonomous agents: multi-step workflows, tool orchestration, long-duration tasks
  • 256,000 token context, 32,000 token output length, 23 languages
  • Requires only 2 A100/H100 GPUs — versus 8 for comparable-power competing models according to Cohere
  • Available on Oracle OCI Generative AI, AWS Bedrock, Azure

Command A Reasoning (August 2025, 111B parameters)

  • Cohere’s first reasoning model
  • Configurable token budget: a low budget produces a fast, economical response; a high budget activates deeper reasoning
  • Available on Oracle OCI, AWS Bedrock
  • No independent benchmark comparing Command A Reasoning to o3 (OpenAI) or Claude 3.7 (Anthropic) was found in sources reviewed

Command A Vision (2025)

  • Multimodal vision + text extension of Command A
  • Available on Oracle OCI Generative AI

Rerank and Embed — on-premises-deployable RAG building blocks

Beyond conversational models, Cohere offers two components that integrate into existing RAG pipelines:

Rerank 3 / Rerank v3.5

  • Reranking model: after an initial broad retrieval, it sorts candidate documents from most to least relevant
  • Pricing: $2.00 per 1,000 searches (1 query + up to 100 documents)
  • Deployable in VPC or on-premises — the document flow never leaves the client’s infrastructure

Embed v3 / v4.0

  • High-performance multilingual embeddings model (encodes up to 128,000 tokens per chunk in v4.0, versus 512 for the previous generation)
  • Pricing: $0.10 per million tokens
  • Referenced in third-party benchmarks such as MTEB among competitive embedders, alongside OpenAI and Voyage

Vocabulary

  • Embedding: numerical representation of text as a vector (list of numbers). Two semantically similar texts produce vectors that are close together in the vector space.
  • Reranking: refinement step in a RAG pipeline. The retriever fetches 50 to 200 candidate documents; the reranker reads them more carefully and retains only the 5 to 10 most relevant before passing them to the generative model.
  • VPC (Virtual Private Cloud): a cloistered cloud environment, accessible only to the organization’s own resources. Data does not transit through the provider’s shared infrastructure.

In plain terms: Rerank and Embed are Cohere’s “silent” components — they don’t write answers, but they make other models’ answers more precise and better sourced. Their on-premises deployment is the decisive argument for sectors where data cannot leave the organization’s perimeter.

Cohere Labs and the Aya project — open-source multilingualism

Cohere Labs (formerly Cohere For AI, renamed April 2025) is a distinct entity from the commercial company Cohere. It is a non-profit research laboratory, founded in June 2022 by Sara Hooker, then led by Marzieh Fadaee since September 2025.

The distinction is fundamental: Cohere Labs produces open-weight models, free, aimed at academic and research use. Cohere (the company) sells paid enterprise models with private deployment. The two entities are not interchangeable.

Aya 23 (May 2024, 8B + 35B parameters) — first Aya model: 23 languages, open weights.

Aya Expanse (October 2024, 8B + 32B parameters) — 101 languages, advanced instruction tuning, cross-lingual transfer. This is the most widely covered multilingual open-source model in the Aya family to date. According to TechCrunch, Cohere claims Aya Expanse outperforms comparable models on multilingual benchmarks [claim reported by press, no independent validation found].

Aya Vision (March 2025) — multimodal vision + text extension, 23 languages, open weights.

The Aya community brings together more than 3,000 researchers in 119 countries, with an explicit focus on languages underrepresented in dominant LLMs.

In plain terms: when talking about Aya, we are talking about Cohere Labs (nonprofit, open source). When talking about Command, we are talking about Cohere (company, paid API). This distinction is critical for evaluating what Cohere actually makes accessible and what remains behind an enterprise paywall.


Positioning — three strategic bets

Bet 1 — Enterprise-only, no competing consumer product

Aidan Gomez states this explicitly: Cohere will not launch a consumer equivalent of ChatGPT. This is not a resource constraint — it is a deliberate choice. The reasoning: an actor that sells LLMs to enterprises while competing with those same enterprises in their end markets loses credibility. Microsoft (OpenAI) sells Copilot to enterprises while offering competing tools; Google integrates Gemini into Workspace. Cohere avoids this conflict.

A direct consequence for the commercial argument: Cohere’s LLM does not accumulate customer usage data to train its own competing products. This is a verifiable non-compete promise.

CriterionCohereOpenAIAnthropic
Consumer productNoYes (ChatGPT)Yes (Claude.ai)
On-premises deploymentYes (native)No (cloud API only)No (cloud API only)
Competition with end customersNoPartial (Microsoft)No
Priority marketB2B enterpriseConsumer + B2BResearchers + B2B

Bet 2 — Data sovereignty and private deployment

Cohere is one of the few players to allow its models to run outside its own cloud infrastructure. Rerank and Embed models can run in the client’s VPC or on their own servers. Command R and Command A are available on sovereign cloud platforms (Oracle OCI, Fujitsu Cloud), not only on the three major hyperscalers.

This architecture responds to a real constraint for regulated sectors: GDPR, HIPAA (healthcare), banking secrecy, defense data. An LLM whose data never leaves the organization’s infrastructure is legally less exposed than a cloud-based shared-API LLM.

The Canadian government materialized this argument by investing 240 M$ in Cohere in December 2024 — an explicit recognition of the company’s role in national AI sovereignty. Partnerships with Saab (Swedish defense), Thales Group, and Hanwha Ocean (Korean naval) illustrate the international reach of this positioning.

In plain terms: Cohere doesn’t just sell a model, it sells the guarantee that the model stays within the organization’s walls. For a hospital, a bank, or a government, this is often the prerequisite for adoption.

Bet 3 — Multilingualism and emerging geographies

Most dominant LLMs are heavily optimized for English. Command A covers 23 enterprise languages; Aya Expanse (via Cohere Labs) covers 101. This differentiator targets several markets underserved by native LLMs: Japan (Fujitsu partnership + Takane LLM), Korea (LG CNS), French-speaking markets, Arabic-speaking markets.

The presence of governments in funding rounds — Canada, and indirectly Japan via Fujitsu — is not anecdotal: it signals that some states see in Cohere an alternative to American providers for building their national AI infrastructure.


Strengths

RAG integrated at the model level Citations in Command R and Command A are produced by the model itself, not by a post-processing wrapper. The model attributes each passage of its response to the exact source chunk. In an enterprise context where response auditability is critical (compliance, legal, medical), this is a differentiating feature.

Computational efficiency Command A requires 2 A100/H100 GPUs for inference according to Cohere — where comparable competing models need 8. Command R7B runs on modest hardware. This reduces total cost of ownership (TCO) for organizations deploying on-premises.

Solid partner ecosystem Oracle (investor + priority distribution), Fujitsu (co-development of Japanese LLM Takane), LG CNS (Korea), Saab, Thales, Dell, RBC (Royal Bank of Canada), Bell Canada, SAP. The geographic coverage — North America, Europe, Asia-Pacific — is consistent with global ambition.

Native multi-cloud Cohere is simultaneously present on AWS Bedrock, Azure AI Studio, Google Cloud Vertex AI, Oracle OCI, and Fujitsu Cloud. Anthropic prioritizes AWS + Google; OpenAI prioritizes Azure. Cohere gives organizations the flexibility to avoid lock-in to a single hyperscaler.


Limitations and areas of attention

Scale and funding versus frontier labs Cohere is valued at 7 Bn$ in September 2025, versus 380 Bn$ for Anthropic and 730 Bn$ for OpenAI in 2026. This scale difference translates directly into the capacity to fund the most expensive training runs and attract top researchers.

Self-declared benchmarks — epistemic position Cohere’s performance claims require careful treatment:

  • Command R+ outperforms GPT-4 Turbo and Claude 3 on RAG (April 2024): from Cohere internal evaluations, no independent published replication found in sources reviewed [not independently replicated].
  • Command R7B first on IFeval, BBH, GPQA, MuSR, MMLU in its size class: Cohere evaluations [not independently replicated].
  • Command A Reasoning versus o3 or Claude 3.7: no independent benchmark found.

This pattern is common in the industry, but it requires distinguishing “Cohere claims that” from “third parties have measured that.”

Lower public visibility The absence of a consumer product reduces name recognition among individual developers and end users. OpenAI and Anthropic build their reputation partly through millions of users who freely test their models. Cohere does not benefit from this distribution and indirect evaluation channel.

Undeclared profitability 240 M$ in annualized revenues (Feb. 2026) against ~1.6 Bn$ raised since 2019, with ~843 employees and significant infrastructure costs. Profitability has not been publicly declared. The announced IPO assumes it will be reached or near before listing.

The sovereignty paradox Cohere sells data sovereignty, but its models are trained on third-party cloud infrastructure (AWS, Oracle, Azure). On-premises deployment concerns inference (the model responding to queries), not training (the process that creates the model). This nuance is rarely highlighted in commercial communications.

RAG positioning — risk of obsolescence In 2024-2026, very long context models (Gemini 2.0 Pro: 1 M tokens, Claude Opus 4.6: 1 M tokens) are absorbing some use cases previously addressed by RAG. Cohere anticipated this risk by pivoting toward agentic with Command A — but the “native RAG” differentiator is gradually losing its exclusivity.

Sara Hooker’s departure Sara Hooker, founding director of Cohere Labs, left the lab in September 2025. She was a recognized scientific figure (Google Brain, co-author of Aya papers). The impact on Cohere Labs’ academic attractiveness remains to be assessed over the next recruitment cohorts.


When Cohere, when another player?

Choosing an enterprise LLM provider depends on specific constraints. Cohere addresses some better than its competitors, and others less well.

ContextRecommendationWhy
Sensitive data that cannot leave the infrastructureCohere (Rerank, Embed, Command A on-premises)Native private deployment, not just cloud API
Need for source citations in responsesCohere Command R/ANative grounding at model level, not post-processing
Languages beyond dominant EnglishCohere Command A (23 languages) or Aya Expanse (101 languages, open)Multilingual coverage exceeding enterprise competitors
Maximize reasoning capabilitiesOpenAI o3, Anthropic Claude 3.7Command A Reasoning without available independent comparative benchmarks
Developer notoriety and adoptionOpenAI GPT-4/5, Anthropic ClaudeLarger ecosystem, more community documentation
Open-source multilingualism, researchAya Expanse (Cohere Labs)Open weights 101 languages, free
Governments and national sovereigntyCohereGovernment investments (Canada), defense partnerships (Saab, Thales)

Heuristic rules

  • If data cannot leave the infrastructure, Cohere is one of the few enterprise players offering a real on-premises architecture for inference. The solution exists today, without workarounds.
  • If the use case is documentary RAG, Command R/A’s native grounding reduces the pipeline development cost — less code to write to handle citations.
  • If advanced reasoning budget is the priority, wait for independent benchmarks on Command A Reasoning before comparing it to o3 or Claude 3.7.
  • Do not confuse Cohere (company, paid models) and Cohere Labs (nonprofit, open-weight Aya) — usage rights, pricing, and license constraints are fundamentally different.

Key takeaways

Cohere occupies a clear strategic niche in the LLM ecosystem: the company that refuses to compete with its customers. This enterprise-only positioning, deployable on-premises, with native grounding and extended multilingualism, addresses real constraints that frontier labs (OpenAI, Anthropic, Google DeepMind) do not prioritize in the same way.

Its strengths are documented and verifiable: the deployment architecture, industrial partnerships, linguistic coverage. Its performance claims, however, remain mostly self-declared and deserve to be treated with caution until independent replication.

The 2024-2026 trajectory shows a company anticipating market evolution — from RAG-first toward agentic — while maintaining its differentiating advantage on sovereignty and private deployment. The announced IPO will be the first real market test of this positioning.

The Cohere / Cohere Labs distinction remains an essential interpretive lens: one sells enterprise solutions; the other produces open-source research on multilingualism. These two entities are not substitutable.