In brief
A dense language model activates all of its parameters for every token it processes. A Mixture of Experts (MoE) model works differently: it maintains a large set of specialized sub-networks — the experts — and activates only a small selection of them for each token. The result: a model can have hundreds of billions of parameters in memory while mobilizing only a few tens of billions at inference time. This decoupling between total capacity and processing cost is what made the MoE paradigm central to LLM architectures since 2021.
In short: imagine a medical practice with 256 specialists. For each patient, an operator (the router) reads the file and calls 2 relevant specialists — not all 256. The whole practice is gigantic, but each consultation only mobilizes a fraction. That is what an MoE model does: DeepSeek-V3 has 671 billion parameters in memory but only mobilizes 37 billion per processed token.
The problem MoE solves
Training a larger language model generally improves its performance. But compute cost grows proportionally to size: doubling the number of parameters roughly doubles the resources required for each forward pass. This proportionality is the main obstacle to scaling.
The MoE intuition predates Transformers. Its modern founding act is a 2017 paper by Noam Shazeer, Geoffrey Hinton, and Jeff Dean: integrating sparse MoE layers into a neural network, where a routing network (the gating network) decides, for each token, which experts to activate. With N experts and only K active at a time (K ≪ N), the model’s capacity scales with N while the compute cost scales only with K. Shazeer et al. report capacity gains exceeding 1000x with minor efficiency losses — on machine translation and language modeling tasks.
Architecture: how it works
In a standard Transformer, each layer contains an attention block and a feed-forward (FFN) block. In an MoE architecture, the FFN block is replaced by a set of experts — typically distinct FFNs — and a router that selects, for each token, K experts to activate.
The router
The router is a lightweight component (often a simple linear layer followed by a softmax) that produces a score vector over all experts for each token. The K experts with the highest scores are activated. This is called top-K routing, with K=1 or K=2 in the vast majority of current implementations.
Switch Transformer (2021) showed that K=1 — a single expert per token — was sufficient to achieve significant gains while simplifying routing and reducing training instability. The authors report pre-training speedups of up to 7x compared to the dense baseline model, and successfully trained a one-trillion-parameter model in low precision — a first.
Mixtral and DeepSeek: MoE in mainstream LLMs
Mixtral 8x7B (Mistral AI, 2024) marked the entry of sparse MoE into accessible LLMs. Each layer has 8 experts; the router activates 2 per token. The model totals 47 billion parameters but mobilizes only 13 billion per token during inference. Mixtral outperforms Llama 2 70B and GPT-3.5 on most evaluated benchmarks at far lower compute cost. Its Apache 2.0 license opened the paradigm to the open-source community.
| Model | Total parameters | Active parameters / token | Experts / layer | Active K |
|---|---|---|---|---|
| Switch Transformer (2021) | 1,000 B | ~1/N | Configurable | 1 |
| Mixtral 8x7B (2024) | 47 B | 13 B | 8 | 2 |
| DeepSeekMoE 16B (2024) | 16 B | ~3 B | 64 (fine) | Variable |
| DeepSeek-V3 (2024) | 671 B | 37 B | 256 (fine) | Variable |
DeepSeekMoE (2024) introduced two innovations on this foundation. First, finer and more numerous experts — smaller individually and activated in larger numbers, enabling more flexible combinations. Second, shared experts that remain always active, absorbing common knowledge across the corpus and reducing redundancy between routed experts. DeepSeek-V3 (671 billion parameters, 37 billion active per token) is, as of late 2024, one of the largest MoE models trained end-to-end.
In short: the ratio between total parameters and active parameters defines the “specialization rate” of an MoE. Mixtral 8x7B activates 28% (13/47B). DeepSeek-V3 activates 5.5% (37/671B). The lower this ratio, the more the model can potentially “store a lot while keeping inference cost low” — provided the GPU memory keeps up (the next pitfall).
What MoE advantages actually hide
The discourse on MoE emphasizes computational efficiency. The reality is more nuanced.
The load balancing problem
Without a corrective mechanism, the router quickly converges toward degenerate behavior: a few experts capture almost all tokens while others rarely train. This routing collapse makes the model as inefficient as a small dense model. To prevent it, MoE architectures introduce an auxiliary balancing loss — an additional constraint that forces the router to distribute tokens more uniformly.
The problem is that this loss interferes with the main model’s gradients and harms expert specialization. ST-MoE (2022) proposed z-loss as a partial solution. DeepSeek-V3 claims to have eliminated all auxiliary loss by dynamically adjusting router biases — an approach that, if confirmed at scale, would resolve one of the paradigm’s fundamental tensions.
Memory: the invisible cost
An MoE model must load all of its experts into memory, even though only K are active at each compute step. The MoE-CAP benchmark documents that MoE models require between 4x and 14x more GPU memory than a dense model of equivalent performance. A model advertised as “more efficient at inference” can thus be more expensive to deploy when memory is the actual constraint — which is the case in the majority of practical configurations.
In short: an MoE saves FLOPs (compute) but consumes massive memory (storage). The operator selected 2 experts out of 256, but the practice must rent a building for 256 specialists. On a GPU with 80 GB VRAM, a 100B-total MoE may not fit when an equivalent 30B dense model fits. VRAM is often the real bottleneck, not compute.
FLOPs don’t measure everything
Dense/MoE comparisons typically rely on FLOPs (floating-point operations). Yet MoE models introduce all-to-all communications between GPUs to route tokens to experts distributed across different devices — a cost absent from FLOP calculations. When measuring actual time per step (which includes both compute and communication), the MoE advantage shrinks compared to FLOP-only comparisons.
The rigidity of top-K
A less-discussed limitation: training a model with a fixed K creates co-dependence among experts. Each expert learns to collaborate with exactly K-1 stable partners. If K is reduced at inference time to gain speed, performance degrades disproportionately. The model has no elasticity in its sparsity degree — a property dense architectures never need to manage.
Do experts actually specialize?
The implicit assumption of MoE is that different experts end up processing distinct content types — some specializing in code, others in mathematical reasoning, others in specific languages. Recent work does find specialization patterns by domain and vocabulary. Semantic probing studies show that the overlap rate between experts varies with the contextual meaning of a word — suggesting genuine semantic sensitivity in the router.
But this specialization is partial, emergent, and uncontrolled. It appears as a byproduct of training rather than as a property guaranteed by the architecture. Work from 2025 also indicates that uniform balancing loss tends to homogenize experts — working against the goal of specialization. The question of how much experts actually encode genuinely distinct competencies remains open.
Decision matrix — MoE vs dense for which use
| Deployment context | Recommendation | Why |
|---|---|---|
| Cloud inference with high-VRAM GPUs (≥ 80 GB × N) | MoE (Mixtral, DeepSeek-V3) | Leverages capacity/compute decoupling, high throughput. |
| On-device inference (mobile, edge, 16-24 GB GPU) | Dense (Llama 3, Mistral 7B) | MoE does not fit in VRAM, equivalent-performance dense is enough. |
| Training with limited compute budget, varied skills | MoE | High total capacity at reasonable FLOPs cost. |
| Training with limited memory budget | Dense | MoE 4-14× more costly in VRAM (MoE-CAP). |
| Production with stable latency constraint < 100 ms | Dense | MoE routing introduces inter-GPU communications (all-to-all), variable latency. |
| Multi-language, multi-domain research with desired specialization | MoE (DeepSeekMoE, fine-grained) | Partial emergent specialization, but usable for scaling. |
Key takeaways
- MoE decouples total model capacity (total parameter count) from per-token compute cost: only a fraction of experts is activated at each pass, via a learned router.
- This architecture enables training very large models at contained compute cost — Mixtral 8x7B activates 13 billion of 47 billion parameters; DeepSeek-V3 activates 37 billion of 671 billion.
- FLOPs advantages do not mechanically translate into deployment advantages: MoE models consume 4 to 14 times more GPU memory than a dense model of equivalent performance, and generate costly inter-GPU communications.
- Load balancing — maintaining an even token distribution across experts — is a structural constraint that interferes with expert specialization. No fully satisfactory solution has yet been validated at scale.
- The fixed top-K rigidity limits inference elasticity: reducing the number of active experts after training degrades performance disproportionately.