In brief
DeepSeek is an artificial intelligence laboratory founded in July 2023 in Hangzhou, China, wholly funded by High-Flyer, a quantitative hedge fund. Unlike most leading laboratories, DeepSeek has no API or cloud revenue model: its mission is research, and all its major models are released as open weights — freely downloadable weights, under the MIT license since January 2025.
In December 2024, DeepSeek published V3, a model with 671 billion parameters and an innovative architecture. In January 2025, R1 became the first large open-weights reasoning model with performance comparable to the best proprietary models on several benchmarks. These two publications had a measurable impact on financial markets and reignited the debate about the true cost of frontier AI development.
Identity card
| Field | Value |
|---|---|
| Founded | 17 July 2023, Hangzhou (Zhejiang, China) |
| Founder and CEO | Liang Wenfeng |
| Majority shareholder | High-Flyer (quant fund, ~84% via holding companies) |
| Headquarters | Hangzhou |
| Flagship models | DeepSeek-V3, DeepSeek-R1 |
| Access | Open weights (MIT since January 2025) |
| Model revenue | Marginal — not the economic driver |
Understanding the context: a lab born from a quant fund
To understand DeepSeek, you first need to understand High-Flyer. Founded in February 2016 by Liang Wenfeng, High-Flyer is a hedge fund that uses quantitative models — statistical algorithms — to trade financial markets. As early as 2021, aware of AI’s potential for high-frequency trading, Liang began buying Nvidia chips in bulk, well before the GPU race became a mainstream topic.
This forward-looking decision had structural consequences: when DeepSeek was created in 2023, High-Flyer already operated a GPU cluster estimated at around 50,000 Hopper chips (H800 and H100), shared between trading operations and AI research. DeepSeek therefore does not need to rent compute — it runs on proprietary infrastructure that is already amortised.
In plain terms: DeepSeek is not a startup seeking a valuation. It is a research lab backed by a fund that needs no immediate returns and already owns the required GPU infrastructure.
A distinct position in the ecosystem
This structure determines DeepSeek’s positioning within the ecosystem:
- Mistral (France) seeks a balance between open-weights releases and a commercial API offering.
- Meta releases Llama to influence the ecosystem and attract engineers, under a conditional licence.
- DeepSeek publishes under MIT, without a premium API, with no economic reason to retain value behind a paywall.
The motivation is explicitly research and scientific prestige, not direct monetisation. All major models (V2, V3, R1) come with technical reports published on arXiv, including training data details, architectural choices, and documented ablations.
Timeline: two years of structural publications
November 2023 — First publications: DeepSeek-Coder, then DeepSeek-LLM 67B, which outperforms LLaMA-2 70B on reasoning, code, mathematics, and Chinese comprehension benchmarks.
January 2024 — DeepSeek-MoE: first research model exploring sparse architecture with shared experts, using only 40% of the compute of an equivalent dense model.
April 2024 — DeepSeek-Math: a 7B model reaching 51.7% on the MATH benchmark (competition level), without any external tooling.
May–June 2024 — DeepSeek-V2: first large public MoE model, introducing the MLA architecture. Followed by Coder-V2, with performance comparable to GPT-4 Turbo on code.
December 2024 — DeepSeek-V3: 671B total parameters, technical report on arXiv (2412.19437). Submitted on 27 December, it introduces several architectural innovations and reports a declared marginal training cost of $5.576 million.
20 January 2025 — DeepSeek-R1: 671B reasoning model, MIT licence, with six distilled variants ranging from 1.5B to 70B. Immediate market impact.
March 2025 — V3-0324: post-training update integrating R1’s RL techniques on the V3 base.
April 2025 — Prover-V2-671B: open-source model specialised in formal proof (Lean 4).
August–December 2025 — Series of updates: V3.1, V3.2-Exp, V3.2, and V3.2-Speciale. This last variant reportedly achieves gold-medal level at IMO, CMO, ICPC World Finals, and IOI 2025 according to DeepSeek — a figure drawn from official documentation, not independently verified.
Early 2026 — Discussion of a $300 million funding round at an estimated $10 billion valuation. In February 2026, Anthropic alleged that DeepSeek used thousands of fraudulent accounts to extract training data from Claude — [UNVERIFIED], a single source (an opposing party), not independently confirmed.
24 April 2026 — DeepSeek-V4 preview, released and open-weighted. The family has two models, both with a one-million-token context window: V4-Pro (1,600 billion total parameters, 49 billion active) and V4-Flash (284 billion total, 13 billion active). Two architectural novelties are announced — a hybrid attention combining Compressed Sparse Attention and Heavily Compressed Attention, and three reasoning effort levels (low / high / max) exposed to the caller.
31 July 2026 — V4-Flash-0731, a production checkpoint oriented towards agents.
13 August 2026 — V4-Pro reaches general availability on the app, the web, and the API.
In short: the ratio that defines DeepSeek did not move between V3 and V4 — it is still an enormous model of which only a small fraction is switched on per token. V3 activated 37 of 671 billion parameters, or 5.5%. V4-Pro activates 49 of 1,600, or 3.1%. The model more than doubled in size, and the cost per token went down.
Architecture: what V3 and R1 changed
DeepSeek-V3: three architectural innovations combined
V3 is a Mixture of Experts (MoE) model: out of 671 billion total parameters, only 37 billion are activated for each token processed. This means that at each operation, the model mobilises roughly 5% of its total capacity — a computational efficiency that distinguishes MoE models from dense ones.
Vocabulary
- MoE (Mixture of Experts): an architecture in which several specialised networks (“experts”) coexist. For each token, a routing mechanism selects the most relevant experts, the only ones performing computations.
- Active parameters: the subset of parameters actually engaged during inference. In a dense model, all parameters are active for every token.
- KV cache: GPU memory space storing intermediate attention results to avoid recomputing them at each generated token.
The technical report (arXiv:2412.19437) documents three architectural innovations:
MLA (Multi-head Latent Attention) — Standard attention maintains a memory cache (KV cache) that grows linearly with context length. MLA compresses this cache via a low-dimensional latent projection, substantially reducing inference memory without measurable quality loss.
DeepSeekMoE — A refined MoE architecture with two categories of experts: “fine-grained” experts (more numerous and more specialised) and “shared” experts (active for every token). This organisation enables load balancing without auxiliary loss functions — a technical simplification relative to most existing MoE models.
FP8 mixed precision — Training in 8-bit precision (versus the standard 16-bit) reduces memory requirements and accelerates computation. The report presents this implementation as the first at scale without measurable degradation — a claim documented in the paper, not independently verified.
In plain terms: V3 combines three simultaneous optimisations — less memory for attention (MLA), less compute per token (MoE), and less memory for training (FP8). The combined effect explains the declared performance per unit of cost.
V3 is trained on 14.8 trillion tokens. Post-training includes SFT (Supervised Fine-Tuning) and RL (Reinforcement Learning). The report states that no rollback or loss spike was observed during training — a declaration not independently verifiable.
DeepSeek-R1: making reasoning emerge through reinforcement
R1 shares the same base as V3 (671B) but with a different objective: producing a model capable of chain-of-thought reasoning — explicitly developing intermediate steps before delivering an answer.
In plain terms: R1 is trained to “think aloud” before answering, like a student showing their working. This strategy improves performance on complex problems that require multiple reasoning steps.
Two variants document the training process (arXiv:2501.12948):
R1-Zero — Trained solely by RL (without prior supervised data), using GRPO (Group Relative Policy Optimization). The model learns to reason through trial and error, guided by no human examples. The report shows that reasoning behaviours — such as reviewing and self-correcting — emerge spontaneously from the reinforcement process alone.
R1 — Full pipeline: cold start data, several RL stages, then a final supervised SFT phase. This variant is the one published under the MIT licence.
Vocabulary
- GRPO (Group Relative Policy Optimization): a variant of the PPO algorithm used for reinforcement-based alignment. Instead of comparing a response to an absolute baseline, GRPO compares responses within the same batch, reducing advantage estimation variance.
- Cold start data: a small high-quality supervised dataset used to initialise RL — prevents the model from starting with excessively random behaviour.
Distilled models: reasoning in smaller formats
Simultaneously with R1, DeepSeek published six distilled models — smaller models trained to imitate R1’s behaviour:
| Distilled model | Base | AIME 2024 | MATH-500 |
|---|---|---|---|
| R1-Distill-Qwen-1.5B | Qwen2.5-Math-1.5B | 28.9% | 83.9% |
| R1-Distill-Qwen-7B | Qwen2.5-Math-7B | 55.5% | 92.8% |
| R1-Distill-Llama-8B | Llama-3.1-8B | 50.4% | 89.1% |
| R1-Distill-Qwen-14B | Qwen2.5-14B | 69.7% | 93.9% |
| R1-Distill-Qwen-32B | Qwen2.5-32B | 72.6% | 94.3% |
| R1-Distill-Llama-70B | Llama-3.3-70B | 70.0% | 94.5% |
Scores self-reported by DeepSeek. The distilled models are themselves open weights.
R1-Distill-Qwen-32B scores 72.6% on AIME 2024, compared to 63.6% for o1-mini according to benchmarks published by DeepSeek — figures replicated by several sources citing the paper, but not independently triangulated.
In plain terms: distillation transfers the reasoning capabilities of a giant model (671B) to models accessible on consumer GPUs. The result is a set of moderately-sized models with unexpectedly strong reasoning performance for their size.
Hugging Face’s open-r1 initiative partially reproduced the scores of distilled models on MATH-500 (±1 to 3 standard deviations across sizes), without reproducing the complete end-to-end R1 training pipeline.
Benchmarks and comparisons
Benchmarks published by DeepSeek for the full R1 model:
- AIME 2024: R1 = 79.8%, OpenAI o1-1217 = 79.2% (source: TechCrunch, DataCamp — reprinting the DeepSeek paper).
- MATH-500: R1 = 97.3%, OpenAI o1-1217 = 96.4% (same sourcing).
- SWE-bench Verified: scores comparable to o1 according to the report.
Important caveat: all these benchmarks are self-reported. The Wall Street Journal reported that on 15 AIME 2024 problems, o1 solves the same problems faster than R1, nuancing the absolute superiority claims. These data come from a single secondary source (WSJ, cited by Wikipedia) and have not been systematically reproduced.
The training cost controversy
The figure and its exact meaning
The V3 report declares 2.788 million H800 GPU-hours for training, equating to a marginal cost of $5.576 million at H800 rental prices. This figure is the cost of the final training run, as explicitly defined in the report itself.
In plain terms: $5.576M is the cost of producing the last version of the model, once all prior work (research, ablations, architecture) is complete. It is not the cost of developing frontier capability.
What this figure excludes
The report itself specifies the exclusions: prior architectural research, ablations, preliminary experiments, infrastructure, and personnel. Independent analyses place the true total cost at a very different level:
- SemiAnalysis estimates the server infrastructure spending of DeepSeek/High-Flyer at approximately $1.6 billion in CapEx, with an annual opex of approximately $944 million.
- SemiAnalysis’s estimated GPU fleet: ~50,000 Hopper GPUs (H100 + H800 + H20) shared between High-Flyer and DeepSeek.
- Martin Vechev (INSAIT, Sofia) stated: “the $6M figure is misleading — it does not represent the cost of developing frontier capability.” Other analyses (Gregory Bufithis, Techstrong.ai) place the total cost between $1.3 and $1.6 billion.
Circular sourcing dependency
The “$5.576M” claim is replicated (at least two official sources: the report and arXiv), but most sources citing this figure point solely to the DeepSeek technical report. No independent verification of the declared cost exists — all press reprints trace back to the same primary source.
In plain terms: the figure of $5.576M is factual within its narrow scope (final run), but using it as a measure of V3’s development cost is incorrect. The debate is real and documented — both camps are sourced.
Open weights and hardware constraints
The MIT licence and its implications
Since January 2025, DeepSeek’s models (V3, R1, and subsequent releases) are published under the MIT licence — the most permissive open-source licence. Unlike Llama (Meta licence with a 700-million-user clause) or Mistral (dual open/proprietary strategy), DeepSeek weights can be integrated into commercial products without restriction or declaration.
The open-weights vs open-source distinction still applies: training data is not published, and the complete training code is not available. Transparency is technical (detailed arXiv reports, downloadable weights) but not total.
Export controls as a technical constraint
DeepSeek trains its models on H800 chips — a version of Nvidia’s H100 specifically designed for export to China, with NVLink bandwidth halved relative to the H100. US export restrictions introduced in 2022 prohibited shipment of H100 and A100 chips to China.
This hardware constraint has two documented effects. First, it pushed DeepSeek to optimise its algorithms to compensate: the combination of MLA + MoE + FP8 can be read as an engineering response to real computational constraints. Second, achieving comparable results with less powerful chips is technically noteworthy, independently of any geopolitical framing.
Impact on the ecosystem
The event of 27 January 2025
The publication of R1 on 20 January 2025 triggered a measurable market reaction. On 27 January 2025, Nvidia’s stock fell 16.9% in a single day, erasing $589 billion in market capitalisation — the largest single-day loss in financial market history according to available data. Hyperscalers (Microsoft, Google, Amazon, Meta) were questioned on the relevance of their GPU investment plans, estimated at $527 billion for 2025–2026.
In practice, no hyperscaler reduced its spending: 2025–2026 investment plans were maintained or accelerated. The logic is that if frontier-quality models are cheaper to produce, inference demand — and therefore GPU demand — increases proportionally.
The redistribution effect within the community
MIT licences and weight publication produced two concrete effects in the ecosystem:
Direct adoption: R1 and its distillations were integrated into third-party products immediately after publication — productivity applications, coding tools, local interfaces. The immediate availability of weights without API access delays explains the speed of adoption.
Community reproduction: Hugging Face’s open-r1 initiative undertook to reproduce R1’s training steps using public datasets such as Numina-Math. This partial reproduction — distilled model scores on MATH-500 confirmed to ±1–3 standard deviations — constitutes the most solid external validation available to date. A complete end-to-end reproduction of the R1 pipeline has not yet been published.
Limitations and unknowns
Benchmark reproducibility: R1 and V3 scores are self-reported. Systematic independent validation does not exist for the full model. Open-r1 partially covers the distilled models.
Training data opacity: as with Llama and Mistral, the exact training corpora are not published. What DeepSeek calls “quality-filtered data” remains unverifiable.
DeepSeek / High-Flyer boundary: the share of the ~50,000 GPU fleet actually allocated to DeepSeek versus High-Flyer’s quantitative trading has not been disclosed. Cost estimates are therefore indirect.
Claude data-scraping allegation [UNVERIFIED]: in February 2026, Anthropic declared that DeepSeek had used thousands of fraudulent accounts to extract data from Claude. This claim has not been confirmed by an independent source — it originates from the opposing party and should be treated with corresponding caution.
2026 trajectory uncertain: Reuters and technical blogs mention a possible V4 in mid-2026, but this information remains speculative at the time of publication.
When to use DeepSeek models
| Context | Model | Why |
|---|---|---|
| Mathematical reasoning, logic, complex code | Full R1 or distilled ≥ 32B | Reasoning performance competitive with o1 on published benchmarks |
| Local deployment on consumer GPU | R1-Distill-Qwen-32B or distilled Llama-70B | Favourable size/performance ratio, unrestricted MIT licence |
| General-purpose production, multilingual, long context | V3 or V3-0324 | Efficient MoE architecture, strong context capacity, 14.8T training tokens |
| Formal proof, competition mathematics | Prover-V2-671B | Lean 4 specialist, declared IMO/CMO performance (not independently verified) |
| Comparison with other open-weights labs | See also Llama 4, Mistral | Heterogeneous benchmarks — compare on the target task, not on AIME alone |
Heuristic rules
- Distilled vs base: if resources are limited, a 32B R1 distilled model offers a better performance/cost ratio than a non-distilled generalist model of the same size for reasoning tasks.
- Check the benchmark before concluding: AIME and MATH-500 measure mathematical reasoning, not general-purpose capabilities. A high score on these metrics says nothing about production code performance or document comprehension.
- MIT licence ≠ full open source: weights are free, training data remains opaque.
- Declared training cost ≠ development cost: using the V3 report figures as a measure of total investment is incorrect — see the Controversy section.
Key takeaways
DeepSeek is an atypical actor in the AI ecosystem: funded by a quantitative fund, with no immediate revenue pressure, systematically publishing as open weights with detailed technical reports. This configuration allowed it to produce V3 and R1 within two years of its founding, with documented architectural innovations (MLA, DeepSeekMoE, GRPO) and a measurable impact on financial markets and the open-weights ecosystem.
Structural limitations remain significant: self-reported benchmarks, opaque training data, controversy over the declared cost. The figure of $5.576M — widely circulated — is factual within its narrow scope (final training run) but does not represent the cost of developing frontier capability. Independent analysts’ estimates of the true total cost sit in a different register.
What DeepSeek has contributed most durably to the ecosystem: the demonstration that optimised MoE architectures can reach top-tier performance with less active compute, and that distilling reasoning into small models is more effective than previously understood.