In brief
Mistral AI is a French startup founded in 2023 by former engineers from Meta and Google DeepMind. In under three years, it has built a complete lineup of language models, from small embedded models to open-weight frontier models. Its signature: the Mixture of Experts (MoE) architecture, which achieves high performance at reduced inference cost, and an open-licensing policy (Apache 2.0) covering the majority of its range. Mistral is today the principal European alternative to the American giants in the race for large language models.
In short: think of a Mixture of Experts as a team of specialised lawyers where, for each incoming case, the front desk automatically selects the two or three most relevant experts. The whole firm has massive expertise (the “total” model), but only a fraction is activated per case (the “active” parameters). Result: the firm’s invoice is computed on the active lawyers, not on the total headcount — exactly the economic model Mistral applies to LLM inference. On Mistral Large 3, 41 billion parameters work per token out of 675 billion available, that is about 6% effective activation.
Model profile
| Field | Value |
|---|---|
| Organisation | Mistral AI (Paris, France) |
| First version | Mistral 7B — September 2023 |
| Type | Frontier lab, startup |
| Access | Open-weight (Apache 2.0) for the majority of models; proprietary API via la-plateforme.mistral.ai |
| Key architecture | Decoder-only Transformer, Sparse Mixture of Experts (SMoE), sliding window attention |
Timeline
Mistral AI entered the competition in September 2023 with a 7 billion-parameter model that surprised the field: Mistral 7B outperformed Llama 2 13B across all benchmarks despite being half the size. The signal was clear — parameter count is not the sole performance factor.
Three months later, in December 2023, Mixtral 8x7B introduced the SMoE architecture into the public lineup. The model totals 46.7 billion parameters, but activates only 12.9 billion per token. Result: performance comparable to GPT-3.5 and Llama 2 70B, with inference speed six times faster than the latter. The MT-Bench score for the instruction-tuned version reached 8.3.
2024 saw the lineup diversify rapidly. Mistral Large (v1) in February targeted the GPT-4 competitive segment, with a 128,000-token context window. Codestral arrived in May — 22 billion parameters, compatible with more than 80 programming languages. In July, Mistral NeMo (12B, co-produced with NVIDIA) added support for eleven languages including Arabic, Chinese and Japanese. Mistral Large 2 (123B) followed the same week, with marked improvements in code, mathematics and reasoning. In September, Pixtral 12B inaugurated vision capabilities in the lineup: a 400-million-parameter vision encoder trained from scratch, paired with the NeMo decoder.
2025 continued on the same trajectory. Mistral Small 3.1 (24B, March) unified multimodal and multilingual support across 21 languages, deployable on a single RTX 4090 GPU or an M-chip Mac with 32 GB of RAM. Mistral Medium 3 (May) positioned itself in the enterprise segment with performance claimed at more than 90% of Claude Sonnet 3.7 across all benchmarks, at a lower price point. In June, Magistral Small and Medium inaugurated chain-of-thought reasoning in the lineup. December 2025 marked a turning point: Mistral Large 3 (675B total, 41B active, granular MoE architecture, 256k-token window, Apache 2.0) reached second place among non-reasoning open-source models on LMArena, and the Ministral 3 family offered dense models in three sizes (3B, 8B, 14B parameters) for embedded deployments — phones, drones, robots.
In March 2026, Mistral Small 4 (119B total, 6B active, 128 experts, 256k window, Apache 2.0) became the first model in the lineup to unify instruction following, configurable reasoning, multimodal capabilities and agentic coding in a single set of weights.
As of 6 September 2026, the documentation presents Mistral Medium 3.5 (mistral-medium-3504) as the most capable model in the lineup: “our frontier-class multimodal model optimized for agentic and coding use cases.” It is commercial, under a Modified MIT licence.
The catalogue has also widened beyond text: Voxtral TTS (voxtral-tts-2603) for speech synthesis with voice cloning, under CC BY-NC 4.0 — therefore non-commercial — and OCR 4.1 (mistral-ocr-4.1), commercial, producing paragraph-level bounding boxes and structural labels. Codestral remains on a Premier licence.
In short: two inversions worth noting, and they run in opposite directions. The first is naming — Mistral’s most capable model is called Medium, not Large. Mistral thus joins Anthropic and Google, where tier names stopped ordering capability in 2026. The second is licensing: the top model is commercial, while Mistral Large 3 and Small 4 are open under Apache 2.0. At Mistral, “the most open” and “the most capable” no longer name the same model — a far more decisive selection fact than the name, and one that neither of the two reveals.
In short: key numbers to remember on the Mistral trajectory. Mistral 7B in September 2023 — 7 billion parameters outperform Llama 2 13B on every benchmark, that is 2× more efficient per parameter. Mixtral 8x7B in December 2023 — 6× faster than Llama 2 70B at equivalent performance thanks to MoE. Mistral Large 3 in December 2025 — entry into the global top 6 of open-source models, second among non-reasoning models. Each generation has confirmed the directing line: maximize efficiency (quality / inference cost ratio) rather than raw size.
Capabilities
The distinctive point of Mistral is its ratio of active parameters to performance. The SMoE architecture — inherited from foundational Mixture of Experts research — allows models like Mixtral 8x7B or Mistral Large 3 to have high parametric capacity while computing only a fraction of the weights at each inference step. The mechanism is straightforward: a routing network selects, for each token, the two to four most relevant experts (specialised feedforward blocks) from N available. The model’s total capacity grows without the cost of a forward pass growing proportionally.
On published benchmarks, Mixtral 8x7B equals or surpasses Llama 2 70B and GPT-3.5 on the majority of tasks, with six times faster inference. Mistral Large 2 (123B) positions itself alongside the frontier models of 2024 for code, mathematics and multilingual reasoning. Mistral Medium 3 claims more than 90% of Claude Sonnet 3.7’s performance at a lower cost ($0.4/M input tokens, $2/M output tokens).
In short: at comparable benchmarks and lower cost, the Mistral trade-off is generally attractive on routine tasks — writing, multilingual translation, informative dialogue. But the differential narrows on complex tasks, where proprietary models (Claude Opus, GPT-5, Gemini Deep Think) remain ahead. Empirical rule: for 80% of standard use cases, Mistral suffices at 30-50% of the competition’s cost; for the 20% of frontier tasks (deep reasoning, critical production code), the gap remains measurable.
Mistral Large 3 (675B/41B active) was ranked #2 among non-reasoning open-source models and #6 globally in OSS on LMArena in December 2025, alongside Meta Llama 3 and Alibaba Qwen3-Omni.
The lineup also covers edge use cases: Ministral 3 (3B to 14B) is designed for on-device deployment on phones, laptops, drones and robots. Mistral Small 3.1 runs on a single consumer GPU at approximately 150 tokens per second.
The open-licensing policy (Apache 2.0 on the majority of models) is a genuine competitive advantage for developers and organisations wishing to deploy without API dependency.
Choosing the right Mistral model
| Context | Recommendation | Why |
|---|---|---|
| Quick prototype, research, academic fine-tuning | Mistral 7B or Mixtral 8x7B (Apache 2.0) | Open weights with free commercial use, mature HuggingFace community, local execution on consumer GPUs. |
| Embedded mobile / edge (smartphone, drone, robot) | Ministral 3B or 8B | Dense models optimized for 32 GB RAM or less. No direct equivalent at OpenAI or Anthropic. |
| Standard enterprise app, multilingual, standard context | Mistral Medium 3 | Claims 90% of Claude Sonnet 3.7 performance at lower cost ($0.4 / $2 per M tokens). |
| Production load, frontier quality, 256k context | Mistral Large 3 (Apache 2.0) | Global OSS top 6, MoE 675B/41B active, 256k-token window. Downloadable and on-premise deployable. |
| Agentic or coding use, maximum quality, commercial licence acceptable | Mistral Medium 3.5 (mistral-medium-3504) | The most capable model in the lineup according to the documentation, despite the “Medium” name. Modified MIT licence — read it before integrating. |
| Speech synthesis with voice cloning | Voxtral TTS | CC BY-NC 4.0 licence: non-commercial use. Check before any client project. |
| Structured document extraction (layout, tables) | OCR 4.1 | Paragraph-level bounding boxes and structural labels. Premier licence, commercial. |
| Multi-language code generation in production | Codestral | 80+ languages covered, but watch the MNPL license — conditional commercial use, check the clauses. |
| Documented RLHF alignment guarantee and independent benchmarks | OpenAI or Anthropic | Mistral publishes less on its safety pipelines and red teaming — fewer external signals on edge behaviors. |
Known limitations
Mistral does not currently have a dense model equivalent to Meta’s 400B+ (Llama 3.1 405B) or first-tier closed models (GPT-4o, Claude 3.5 Sonnet, Gemini Ultra). Mistral Large 3 reaches 675B total parameters, but only 41B are active per token — direct comparison with a dense model of the same size does not apply. MoE efficiency is real, but the memorisation capacity and generalisation ability of a very large dense model remains an open question.
The lineup’s multimodal integration is more recent than competitors’. Pixtral 12B (2024) is Mistral’s first vision model, while OpenAI and Google had vision capabilities since 2023. The maturity of the vision encoder on complex tasks (spatial reasoning, document understanding) remains to be independently assessed beyond official announcements.
The third-party ecosystem — plugins, integrations, community tools — is significantly smaller than OpenAI’s or Google’s. The number of documented deployments and native integrations in professional tools (IDEs, no-code platforms, CRMs) is below the American competition.
The licensing policy is not uniform. Codestral uses the Mistral AI Non-Production License, which prohibits commercial use without a specific agreement. For an organisation seeking a uniform licence across the full lineup, this heterogeneity requires case-by-case verification.
Finally, public transparency on alignment methods (RLHF, red teaming, safety evaluations) is lower than Anthropic’s or OpenAI’s, making independent assessment of model behaviour in edge cases more difficult.
Key takeaways
- Mistral AI has built in under three years a model lineup that competes on multiple segments with American players, in open-weight form and at lower cost.
- The Mixture of Experts architecture is central to this strategy: it maximises parametric capacity without linearly increasing inference cost.
- The lineup spans from edge (Ministral 3B) to frontier (Mistral Large 3), and includes multimodal, code, reasoning and multilingual capabilities.
- The Apache 2.0 policy on the majority of models is a strong differentiator for use cases requiring local deployment or freedom from API vendor dependency.
- Limitations are real: narrower third-party ecosystem, more recent multimodal capabilities, heterogeneous licences.