In brief
The word “open source” applied to LLMs covers very different realities. Most models presented as open only publish their weights — not their training data, sometimes not their code, and almost never under a truly free licence. In 2025, the performance gap with proprietary models has narrowed considerably, but the debate over what “openness” implies remains very much alive: legally, politically, and on safety grounds.
In short: “open source” for an LLM is like “homemade” for a restaurant dish. The word suggests total transparency (recipe, ingredients, suppliers); in practice, you sometimes only get the dish on the plate (the weights), sometimes also the written recipe (the code), rarely the exact list of ingredients (the data), and almost never the producer’s signature.
The vocabulary confusion that changes everything
When Meta announces that Llama is “open source”, or when DeepSeek publishes R1 under an MIT licence, the word “open” does not mean the same thing in both cases. This confusion is not trivial: it determines what a developer can do with the model, what a company can audit, and what a regulator can demand.
Three levels must be distinguished:
Open weights: the model weights are published. They can be downloaded, run locally, sometimes modified. But the training data remains opaque, the licence may be restrictive, and the training code is not always available. This is the case for Llama 4 (Meta) and Gemma (Google).
Partially open source: weights + code + description of the data, under a permissive licence. Mistral and Qwen fall into this category. The pipeline can be reproduced, without access to the data themselves.
Fully open: weights, code and training data published in full. Only a few academic models reach this level: OLMo (Allen Institute for AI), Pythia (EleutherAI), the LLM360 models. These are exceptions.
In short: three floors of openness, not two. Open weights = you have the dish. Partially open source = you have the dish plus the recipe, but without the supplier list. Fully open = you have everything, recipe and suppliers included. The press systematically conflates the three under the word “open source”.
The Open Source Initiative (OSI) attempted to formalise this with its Open Source AI Definition (OSAID v1.0), published in October 2024. Its central position: it does not require publication of the training data themselves, only a description sufficient to reconstruct an equivalent system. This concession to economic realism — frontier datasets cost billions — provoked objections, notably from the Software Freedom Conservancy, which sees it as a weakening of the foundational principles of free software.
Result: the only models validated by the OSI are Pythia, OLMo, Amber, CrystalCoder and T5. Llama 2, Grok, Phi-2 and Mixtral fail the test.
A heterogeneous licence landscape
The licence determines what you are allowed to do, not just what is technically accessible.
| Model | Licence | Effective freedom |
|---|---|---|
| Llama 4 (Meta) | Llama Community Licence | Forbidden for organisations with >700M monthly active users; forbidden to train competing models |
| Mistral Large 3 | Apache 2.0 | Free commercial use |
| Qwen3 (Alibaba) | Apache 2.0 | Free commercial use, no user cap |
| Gemma (Google) | Gemma Terms of Use | Commercial use with restrictions |
| DeepSeek V3/R1 | MIT | Almost no restrictions |
| Falcon (TII) | Apache 2.0 | Free commercial use |
The Llama licence deserves attention. Meta positions itself as a champion of open AI, but its licence excludes precisely the actors who could use Llama to compete with it directly. This is not open source in the OSI sense: it is an ecosystem strategy — making the model accessible to individual developers and SMEs, while protecting Meta’s commercial territory.
The performance gap is closing
For a long time, choosing between open source and proprietary was also a choice about quality. That is no longer straightforward.
The MMLU gap between the best open-source model and the best proprietary model fell from 17.5 percentage points to 0.3 points in the space of a year. On mathematical benchmarks, Kimi K2.5 reaches 96% on AIME 2025. On code, the best open-source models touch 90% on LiveCodeBench.
The triggering event was the release of DeepSeek-R1 in January 2025. This model, published under an MIT licence, rivalled GPT-4o on several benchmarks at significantly lower compute cost. On the announcement, the Nasdaq fell 3.1%. DeepSeek demonstrated that rigorous algorithmic engineering can compensate for limited access to high-end GPUs — which weakens the thesis that raw compute guarantees proprietary superiority.
The cost advantage remains real: $0.83 per million tokens on average for open source, versus $6.03 for proprietary — an 86% saving. On optimised infrastructure, open-source model latency can reach 3,000 tokens per second, compared to 600 for proprietary APIs.
A gap nevertheless persists. According to whatllm.org, the best open-source model (MiniMax-M2, score 61) remains slightly below the best proprietary model (GPT-5, score 68). The methodology of this ranking is not fully detailed.
What open source makes possible
Local deployment and data sovereignty. An open-weights model can run on internal servers, with no data transmitted to an external API. This is the central argument for finance, healthcare, defence and public administrations. Several countries are building “sovereign AI” strategies on this basis — the EU being a visible example.
Unrestricted fine-tuning. Adapting a model to a specific business domain requires access to the weights. On a proprietary model, fine-tuning tools are limited, controlled by the operator, and dependent on its roadmap. On an open model, adaptation is total.
Auditability. Inspecting the weights makes it possible to detect undocumented behaviours, verify regulatory compliance, and identify potential backdoors. This argument carries weight in regulated sectors.
No vendor dependency. No lock-in to a single supplier. Portability across clouds. No API fees at high volumes.
What open source complicates
Dual use. A published model can be downloaded, modified and redeployed without the original safety filters. Protections built in by RLHF can be removed by fine-tuning. Cited uses include: malware generation, industrial phishing, disinformation, assistance with biological weapon fabrication. Empirical data on how easy these operations actually are remain poorly documented.
Irreversibility. Once published, a model cannot be “recalled”. Unlike a proprietary API whose access can be revoked, an open model is in the wild permanently. OpenAI and Anthropic argue that proprietary models allow post-deployment control — monitoring, guardrail updates, access revocation — that is structurally impossible in open source.
Diffuse responsibility. The law has not resolved the question of the original developer’s liability for malicious uses of a model modified by a third party. The comparison with free software is contested: the potential harm of an LLM is judged to be different from a software library.
Paradoxical concentration. Open source does not eliminate concentration — it shifts it. Publishing a frontier model remains out of reach for any actor without massive compute access. The main publishers are Meta, Alibaba, Google and Microsoft. Truly open academic models (OLMo, Pythia) are produced by resource-constrained actors and do not rival frontier models in performance. Open source guarantees freedom of use, not the capacity to innovate at the highest level.
The EU AI Act and the exemption case
The EU AI Act (applicable to GPAI models from 2 August 2025) includes a partial exemption for open-source models. Conditions: genuinely free licence, no direct monetisation, weights, architecture and usage publicly available.
There is an exception to the exception. Models presenting “systemic risks” do not benefit from the exemption, even if open source. The definition of “systemic risk” in the regulatory text remains vague: no quantitative public criteria have been published by the European Commission. A large open-source model could theoretically be subject to the same obligations as a proprietary model.
One obligation applies to all, regardless of status: compliance with European copyright law across the entire model lifecycle. Models already on the market before 2 August 2025 have a deadline until 2 August 2027.
Hugging Face as an observatory
Hugging Face has become the central infrastructure of open-source LLMs: 13 million users, 2 million public models, 500,000 datasets in 2025. The figures reveal internal concentration: 0.01% of models account for 49.6% of downloads.
The geography of usage is instructive. China overtakes the United States in monthly downloads (approximately 41%). Qwen (Alibaba) counts 113,000 derivative models on the platform, far ahead of Llama (27,000) and DeepSeek (6,000). More than 30% of Fortune 500 companies have verified accounts on Hugging Face.
These data illustrate the geopolitical paradox: the global open-source ecosystem is partly structured around models produced in China, under permissive licences, downloaded en masse by developers worldwide. American export controls on GPUs are partially bypassed by algorithmic efficiency — which is precisely what DeepSeek demonstrated.
Which model to choose by need
| Context | Recommendation | Why |
|---|---|---|
| Confidential customer data, on-prem legal constraint | Open weights (Llama, Mistral) self-hosted | No data transit to a third party; simplified GDPR/HIPAA compliance. |
| Massive request volume, critical marginal cost | Open weights + optimized inference (vLLM, TGI, llama.cpp) | Cost reduced to GPU + electricity, no vendor margin. Break-even vs API ~10⁶ requests/month depending on model size. |
| Time-to-market, team without internal MLOps | Proprietary API (Claude, GPT, Gemini) | Zero ops, frontier performance, enterprise support. Visible cost but easy to model. |
| Regulatory audit or reproducible research | Fully open (OLMo, Pythia, LLM360) | Only level allowing tracing biases, verifying benchmark decontamination, or auditing for EU AI Act. |
| Strong domain specialization (legal, medical, financial) | Open weights + fine-tuning on domain data | Total control of the adaptation process; proprietary APIs limit fine-tuning of frontier models. |
Key takeaways
- “Open source” for LLMs is a term that covers at least three distinct realities: open weights, partially open source, and fully open. Only a few academic models fully satisfy the OSI criteria.
- The licence determines actual usage. Apache 2.0 and MIT offer maximum freedom. The Llama Community Licence protects Meta’s commercial interests despite its “open” presentation.
- The performance gap with proprietary models has narrowed dramatically in 2025. Cost remains the most solid competitive advantage of open source.
- Concrete advantages: data sovereignty, fine-tuning without dependency, auditability, no vendor lock-in.
- Real risks: dual use without guardrails, irreversibility of publication, unresolved legal liability.
- The EU AI Act includes an exemption for open source, but models with “systemic risk” do not benefit — and this category remains to be precisely defined.
- Structural concentration persists: only actors with frontier compute can publish frontier models. Open source shifts concentration; it does not eliminate it.