In brief

In February 2023, Meta released Llama 1. The stated goal was modest: to provide the academic community with a quality language model for research. A few days after its release, the model weights leaked online. What could have been a security incident became the catalyst for a quiet revolution: thousands of developers began experimenting, fine-tuning, and deploying Llama on their own machines.

Since then, Meta has methodically expanded the family. Each generation has pushed the boundaries of what an accessible model can do: longer context, multimodal capabilities, more efficient architecture. With Llama 4, released in April 2025, Meta crossed the billion-token context threshold and adopted a Mixture of Experts architecture. Llama is no longer simply a research model — it has become the underlying infrastructure for a large part of the global open-weights AI ecosystem.

In short: Llama plays the role in the AI ecosystem that Linux played in operating systems in the early 2000s. Not the most polished, not the most marketed — but distributed for free with its weights, modifiable, deployable on your own machine. Where GPT-4 and Claude remain locked behind a paid API, Llama lives in thousands of local configurations, from laptops to clusters. This omnipresence has become a de facto infrastructure that Meta has methodically maintained.


Identity card

FieldValue
OrganizationMeta AI
First versionFebruary 2023 (Llama 1)
TypeFamily of language models (text, then multimodal)
AccessOpen weights (free download under Meta license)
Available sizesFrom 1B to ~400B parameters depending on version

History

Llama 1 (February 2023) marks Meta’s first release in the large language model space. Available in 7B, 13B, 33B, and 65B parameters, it is trained exclusively on public data — CommonCrawl, GitHub, Wikipedia, ArXiv. Its license is restricted to research, but its rapid leak makes it the reference model for the nascent open-source community.

Llama 2 (July 2023) introduces two decisive changes: a conditionally open commercial license, and the first public documentation of a RLHF (Reinforcement Learning from Human Feedback) process from Meta. The training data volume reaches 2 trillion tokens. Chat variants (instruct-tuned) appear, with sizes up to 70B.

Llama 3 (April 2024) represents a qualitative leap. Training covers 15 trillion tokens — seven times more than Llama 2. Code support is multiplied by four. The architecture adopts Grouped-Query Attention (GQA) across all models, improving inference efficiency. Initial sizes are 8B and 70B, before Llama 3.1 adds in July 2024 a flagship model of 405B parameters — then the largest available open-weights model — with an extended context window of 128,000 tokens.

Llama 3.2 (September 2024) marks the family’s entry into multimodality. Two vision variants (11B and 90B) process images and text jointly, trained on six billion image-text pairs. Two lightweight text variants (1B and 3B) target mobile devices and edge deployments with a 128K-token context.

Llama 3.3 (December 2024) is a single 70B model focused on efficiency: its performance approaches that of Llama 3.1 405B according to Meta, at an inference cost approximately five times lower than Llama 3.1 70B.

Llama 4 (April 2025) moves away from dense architecture in favor of Mixture of Experts. Scout (17B active parameters out of 109B total, 16 experts) features a context window of 10 million tokens — a record for open-weights at the time of release — and can run on a single H100 GPU with quantization. Maverick (17B active out of ~400B total, 128 experts) targets maximum performance. A third model, Behemoth (~2,000B total), was announced but not yet released at the time of publication.

In short: keep three scale jumps in mind. Training data: 1 trillion (Llama 1) → 2 trillion (Llama 2) → 15 trillion (Llama 3) — multiplied by 15 in two years. Maximum context: 2,048 (Llama 1) → 128,000 (Llama 3.1) → 10 million (Llama 4 Scout) — that is ×5,000 on the window. Architecture: dense models up to Llama 3.3 → MoE generalized on Llama 4. Each generation has narrowed the gap with proprietary models without changing the economic model: weights freely downloadable.

2026 — the series stops, another one starts

There was no Llama 5. On 8 April 2026, Meta releases Muse Spark, the first model from Meta Superintelligence Labs and the first of a series named Muse — a natively multimodal reasoning model, with tool use, visual chain of thought and multi-agent orchestration. Meta describes it as deliberately small and fast, each generation validating the previous one before going bigger. It powers the Meta AI app and website, then WhatsApp, Instagram, Facebook, Messenger and the AI glasses.

Muse Spark is not open-weight. That is the break. After three years of downloadable weights, Meta’s flagship model becomes a service again.

On 10 August 2026, Meta partly walks it back with Muse Glimmer, a 30-billion-parameter dense model distilled from Muse Spark, with a dedicated perception encoder, released under Apache 2.0. It is designed for local agentic work: multi-step reasoning, tool use, multimodal understanding and failure recovery, all runnable on consumer hardware with no network access.

In short: the “open the weights to become the infrastructure” bet was not abandoned, it was recut. The top of the range closes, the distilled model opens. Meta keeps what gives the competitive advantage and publishes what keeps the ecosystem alive — and what runs on the reader’s machine, not in its datacenter.

Capabilities

Benchmarks (Llama 4, April 2025). On MMMU (multimodal reasoning), Maverick reaches 73.4% versus 69.1% for GPT-4o and 71.7% for Gemini 2.0 Flash. On MathVista, Maverick scores 73.7% versus 63.8% for GPT-4o. Scout achieves 74.3% on MMLU Pro and 57.2% on GPQA Diamond. These figures are published by Meta and had not been subjected to systematic independent validation at the time of release.

Long context. Scout’s 10 million-token window represents a paradigm shift for open-weights: it becomes possible to inject entire corpora into context without resorting to RAG.

Community ecosystem. Llama owes part of its practical relevance to the infrastructure built around it. llama.cpp (ggml-org) enables inference in C/C++ on CPU and GPU, becoming the de facto standard for local execution. The GGUF format (GPT-Generated Unified Format) standardizes the distribution of quantized models, with levels from 1.5 to 8 bits. Ollama simplifies local deployment to a single command. Hugging Face centralizes distribution and conversion tools. For fine-tuning, LoRA and QLoRA allow adapting a model to a specific task with limited resources; frameworks such as Unsloth and LlamaFactory have industrialized this pipeline.

In short: three projects in the Llama ecosystem matter as much as the models themselves. llama.cpp lets you run Llama on a MacBook or a GPU-less server. Ollama simplifies all of this into a single command (ollama run llama3). HuggingFace centralizes the thousands of fine-tuned variants the community produces every month. Without this free tooling built around them, the Llama weights would just be cumbersome binary files. With it, they become a real production infrastructure.

Known limitations

The open weights / open source distinction is not cosmetic. An “open source” model in the sense of the Open Source Initiative (OSI) implies the availability of training code, data, and a license without usage restrictions. Llama is none of these: only the weights are published, training data is not disclosed, and the license imposes conditions. The OSI has explicitly asked Meta to stop using the term “open source” for Llama 2 and 3.

The 700 million users clause is the most visible restriction: any entity exceeding this monthly active user threshold must obtain a separate commercial license from Meta. This clause directly targets major competitors (Google, Microsoft) and signals that Meta’s “generosity” is strategically bounded.

Alignment is less documented than in proprietary models. Meta’s RLHF is described in technical papers but remains less intensive than that of models like GPT-4 or Claude — no quantitative comparison has been published. Base (non-instruct) versions are deliberately unfiltered, which is useful for research but places greater responsibility on deployers.

Hallucinations remain an unsolved problem. Llama 4 uses DPO (Direct Preference Optimization) to improve answer accuracy, but Meta has not published measured hallucination rates. Direct comparisons with other models are difficult to standardize due to heterogeneous metrics across sources.

Benchmarks are self-reported. Performance figures published at Llama launches are produced by Meta. The absence of systematic independent evaluation at the time of releases makes objective comparison with competing models difficult.

When to use Llama over a proprietary model

ContextRecommendationWhy
On-premise deployment for sovereignty or compliance reasonsLlama 3.x or 4Freely downloadable weights, local execution without external API calls — required in regulated sectors (health, defense, certain public markets).
Domain fine-tuning on private datasetLlama 3.x via QLoRAFine-tuning possible on a single consumer GPU (24 GB VRAM). With a closed proprietary model, this adaptation is not possible.
Reproducible academic researchLlama (weights + technical paper)Stable weights are citable and experiments reproducible. GPT-4o or Gemini change silently and invalidate longitudinal comparisons.
High-volume consumer app, constrained API budgetSelf-hosted Llama or specialized cloudAt inference, cost per million tokens drops with size (7B to 70B models) — often 5-10× cheaper than OpenAI or Anthropic at comparable quality.
Application requiring the strongest alignment guaranteeOpenAI or AnthropicLlama RLHF is less documented and less intensive than GPT-4 or Claude. Base unfiltered versions expose more edge behaviors.
Service serving more than 700 million active usersOpenAI or AnthropicLlama license clause requires separate commercial negotiation beyond this threshold — structural friction for Meta’s large competitors.

Key takeaways

  • Llama has changed the structure of the LLM ecosystem: before it, high-performing models were the exclusive domain of major labs with paid API access.
  • Each generation brings open-weights performance closer to proprietary models, while reducing inference costs.
  • Llama 4 Scout, running on a single H100 GPU with a 10 million-token context, illustrates how far this dynamic can go.
  • What Meta publishes is open weights, not open source: training data remains opaque, training code is not published, and the license is conditional.
  • This distinction determines what the community can actually verify, reproduce, and modify.