In brief
In April 2023, Alibaba Cloud presented Tongyi Qianwen — in Mandarin 通义千问, “universal answer to a thousand questions” — first in a closed beta on its own services. Less than three years later, the Qwen family had become the world’s top producer of derivative models on HuggingFace, ahead of Meta and its Llama: 113,000 models directly derived from Qwen, representing 69% of new fine-tunes uploaded to the platform in February 2026. [triangulated — HuggingFace State of Open Source Spring 2026 + Xinhua January 2026 + ATOM Report arXiv 2604.07190]
This trajectory results from a coherent strategy: releasing open-weight models under the Apache 2.0 license, covering the full size range (from 0.6 billion to 235 billion parameters), and integrating into a single model capabilities that competitors require multiple separate systems to deliver — reasoning, vision, code, audio.
Three entities not to confuse
Before diving into the models, a clarification is needed: several distinct organizations are involved in the Qwen production chain, and confusing them leads to strategic misreadings.
Organizational vocabulary
- Alibaba Group : the publicly traded holding company (NYSE: BABA, HKEX: 9988), founded in 1999 by Jack Ma in Hangzhou. The parent company.
- Alibaba Cloud (Aliyun) : Alibaba’s cloud computing division, launched in 2009. The commercial entity that hosts, distributes, and monetizes Qwen models via API.
- DAMO Academy (Academy for Discovery, Adventure, Momentum and Outlook) : Alibaba’s fundamental research laboratory established in 2017, with offices in Hangzhou, Beijing, San Mateo, Singapore, and Moscow. Its mandate spans broad domains: IoT, fintech, quantum computing, AI. DAMO was the initial research entity before the reorganization. [triangulated — Alibaba Group official + MIT Technology Review 2018]
- Tongyi Laboratory (通义实验室, renamed Tongyi Large Model Business Unit) : the operational unit directly responsible for Qwen development since Alibaba’s internal restructuring. In March 2026, a new consolidation created “Alibaba Token Hub” under CEO Eddie Wu, grouping all AI operations. [triangulated — Wikipedia Qwen + Xinhua]
The practical relationship is: Tongyi Lab develops the models, Alibaba Cloud commercializes them via DashScope (API) and Model Studio (SaaS interface), and Alibaba Group drives strategy and investments. Qwen is not an isolated research project — it is the central axis of Alibaba Cloud’s competition against AWS, Azure, and Google Cloud in the cloud AI market.
In plain terms : when you read “Alibaba releases a new model,” the team that produced it is Tongyi Lab, the platform distributing it is Alibaba Cloud, and the group bearing the cost is Alibaba Group. Three entities, one visible product.
Identity card
| Field | Value |
|---|---|
| Organization | Tongyi Laboratory / Alibaba Cloud |
| First public release | September 2023 (Qwen 1, general public) |
| Type | Multi-architecture family (dense + MoE) |
| Access | Apache 2.0 open-weight (public models) + proprietary API (Qwen3-Max, Qwen3.5-Plus) |
| Available sizes | From 0.6B to 235B parameters (Qwen3, April 2025) |
| Languages | 119 languages (Qwen3, 36-trillion-token corpus) |
| License | Apache 2.0 (open-weight since Qwen3, April 2025) |
History: from closed beta to ecosystem dominance (2023–2026)
2023 — Regulated first steps
April 2023: Tongyi Qianwen is presented at a press event in Beijing, deployed in closed beta on DingTalk (Alibaba’s professional messaging) and Tmall Genie (consumer voice assistant). Access is restricted to enterprise clients. [triangulated — SiliconAngle + Wikipedia + Tongyi Qianwen AI Wiki]
The beta phase is not trivial: in China, any public-facing conversational AI service must obtain approval from the Cyberspace Administration of China (CAC). Tongyi Qianwen receives this approval in September 2023 and simultaneously opens to the Chinese general public. The first weights are released under the proprietary Tongyi Qianwen license: 7B, 14B, 72B, and 1.8B parameters. [triangulated — SiliconAngle + Wikipedia + Skywork.ai]
2024 — Multimodal build-up
June 2024 — Qwen2: mixed dense and MoE architectures, with the first serious multimodal variants (Qwen-Audio, Qwen-VL). Published benchmarks indicate performance exceeding Llama 2 70B on math and code tasks. The license partially shifts to Apache 2.0 for models below 72B. [replicated — Wikipedia + arXiv Qwen2 Technical Report]
September 2024 — Qwen2.5: the pre-training corpus grows from 7 to 18 trillion tokens. Qwen2.5-72B-Instruct claims performance comparable to Llama 3.1-405B — a model five times larger — on general reasoning benchmarks. Six sizes of Qwen2.5-Coder (code, 92 programming languages) and three sizes of Qwen2.5-VL (vision-language: 3B, 7B, 72B) are launched simultaneously. [triangulated — arXiv 2412.15115 + HuggingFace papers + Qwen readthedocs]
2024–2025 — The emergence of reasoning
November 2024 — QwQ-32B-Preview: Qwen’s first pure reasoning model, using reinforcement learning with outcome-based rewards. The approach draws directly from DeepSeek-R1, released shortly after. [replicated — GitHub QwenLM/QwQ + VentureBeat]
March 2025 — QwQ-32B (final version): 32 billion parameters. Published benchmarks show performance close to DeepSeek-R1 (671B total, 37B active) on AIME and LiveCodeBench, and superior to OpenAI o1-mini. Open-weight under Apache 2.0. [triangulated — GitHub QwenLM/QwQ + BDTechTalks + HuggingFace QwQ-32B card]
In plain terms : QwQ-32B demonstrates that a 32-billion-parameter model, trained intelligently, can match on certain reasoning tasks a model twenty times larger. This is the same signal as DeepSeek-R1 a few weeks earlier: raw size is not the only lever.
April 2025 — Qwen3, the pivotal release
Qwen3 is the most structurally significant release in the family. Published on April 28–29, 2025, it introduces several major changes simultaneously: [triangulated — qwenlm.github.io + arXiv 2505.09388 + Alibaba Cloud blog]
- Two architectures: six dense models (0.6B, 1.7B, 4B, 8B, 14B, 32B) and two MoE (30B-A3B and 235B-A22B).
- Universal Apache 2.0 for all open-weight — end of licensing heterogeneity.
- 36-trillion-token corpus across 119 languages (versus 18T for Qwen2.5).
- Integrated reasoning: thinking (chained reasoning) and non-thinking (direct response) modes are available in the same model, without model switching.
This last point deserves attention: whereas Mistral released separate Magistral models for reasoning, and OpenAI distinguishes o1/o3 from GPT-4o, Qwen3 integrates both modes in a single weight file. [replicated — qwenlm.github.io + Alibaba Cloud blog]
2026 — Consolidation and new series
Qwen3.5 (Qwen3.5-Plus, proprietary API) and Qwen3-Coder-Next appeared in February 2026. Qwen3.6 was being deployed in April 2026. [replicated — Wikipedia Qwen + GitHub QwenLM/Qwen3.6, single source]
June 2026 — Qwen3.7-Plus: a multimodal agentic model unifying vision and language, presented as an upgrade to vision-language capabilities with no loss on code, tool use, and productivity tasks. [single source — Alibaba Cloud Community]
3 August 2026 — Qwen3.8-Max: the largest model the family has ever released — 2,400 billion parameters, a one-million-token context window, multimodal. Alibaba reports fifth place in Text Arena and second in Vision Arena. The model is first accessible through the API on Alibaba Cloud Model Studio, with weights announced for later. [single source — Alibaba Cloud, press release and qwen3.8-2.4t-a95b model page]
Qwen3.8-Flash-Next, a multimodal MoE model, is presented by Alibaba as a preview of the architecture chosen for Qwen4.
In short: in sixteen months the family went from Qwen3 (235 billion parameters at the top) to Qwen3.8-Max (2,400 billion), a factor of ten. But the flagship model has moved back behind a proprietary API, where Qwen3 had made universal Apache 2.0 its argument. The weights of Qwen3.8-Max are announced, not published — that is a promise, not yet a fact.
The branches of the family
The Qwen family is not a single model but a multi-branch ecosystem. Understanding each branch’s role avoids confusion.
Dense language models (LLM base)
The six dense Qwen3 models — 0.6B, 1.7B, 4B, 8B, 14B, 32B — cover an exceptional range. The 0.6B is designed for edge devices (smartphones, microcontrollers). The 32B targets professional GPU servers. All are Apache 2.0 and available on HuggingFace and ModelScope. [triangulated — qwenlm.github.io + arXiv 2505.09388 + bestcodes.dev]
MoE models (Mixture of Experts)
Vocabulary
- MoE (Mixture of Experts): an architecture where the model is divided into specialized sub-networks (“experts”). For each token processed, only a small subset of experts is activated. Result: the total model size is large (many available parameters) but the per-token cost remains that of a much smaller model.
- Active parameters: the number of parameters actually used during token computation. In a MoE, this figure is well below the total.
Qwen3-30B-A3B: 30 billion total parameters, 3 billion active per token. 128 experts, 8 activated per inference. The production-ready MoE: efficient to deploy, performant for its activated size.
Qwen3-235B-A22B: the flagship MoE. 235 billion total parameters, 22 billion active. Benchmarks published by the Qwen team compare it favorably to DeepSeek-R1, GPT-o1, and Gemini-2.5-Pro on code, mathematics, and general reasoning. The ATOM Report (arXiv 2604.07190) provides partial third-party validation — but not systematic coverage of all claims. [official source for direct comparisons; replicated ATOM Report]
In plain terms : Qwen3-235B-A22B has 235B parameters in total, but only “thinks” with 22B at any given moment. This makes it less costly to run than a dense 235B model, while potentially holding a richer “knowledge library.”
QwQ — the reasoning branch (Qwen2 only)
QwQ-32B is the dedicated chain-of-thought reasoning model for the Qwen2 generation. Since Qwen3, reasoning is directly integrated into standard models — the QwQ branch is specific to the Qwen2→Qwen3 transition and will not continue as an independent family.
Qwen2.5-VL and Qwen3-VL — vision-language
Qwen2.5-VL (January 2025): processing of images, documents, charts, and videos up to one hour long. Available in 3B, 7B, and 72B. Capable of visual agents (computer use) — understanding a computer or phone interface. [triangulated — arXiv 2502.13923 + HuggingFace Qwen2.5-VL + Hackster.io]
Qwen3-VL: the multimodal version of Qwen3, being deployed mid-2025. [replicated — GitHub QwenLM/Qwen3-VL]
Qwen2.5-Coder — code specialist
Six sizes (from 0.5B to 32B), 92 programming languages. The 32B Instruct version claims performance comparable to GPT-4o on EvalPlus, LiveCodeBench, and BigCodeBench, with an Aider score of 73.7 (versus 72.9 for GPT-4o). [triangulated — arXiv 2409.12186 + Ollama library + HuggingFace Qwen2.5-Coder card]
Qwen2.5-Omni — full multimodal
Announced in March 2026: text, images, video, and audio as input; text and audio generation as output. [replicated — Wikipedia Qwen + CNBC 2026]
Strategic positioning
Apache 2.0 open-weight vs proprietary
Since Qwen3 (April 2025), all open-weight models are under Apache 2.0 license — free commercial use with no audience-size restrictions. This converges with Meta’s Llama 4 strategy. Proprietary versions (Qwen3-Max, Qwen3.5-Plus) remain API-only via DashScope. [triangulated — GitHub Qwen LICENSE + HuggingFace Qwen2.5-72B-Instruct + qwenimage.art blog]
In plain terms : Qwen3’s rule is simple — if the model is open-weight, it is Apache 2.0. If it is not Apache 2.0, it is API-only. No licensing ambiguity as in early generations.
ModelScope vs HuggingFace
ModelScope (modelscope.cn) is Alibaba’s open-model platform, positioned as an alternative to HuggingFace for developers in China (where access to HuggingFace can be restricted). Qwen models are available there in priority or simultaneously with HuggingFace. Integration with Alibaba Cloud infrastructure (PAI-EAS, DashScope) is native. [replicated — Alibaba Cloud documentation + GitHub QwenLM]
Alibaba Cloud enterprise infrastructure
The commercial triplet: DashScope (OpenAI-compatible API), Model Studio (SaaS management interface), PAI-EAS (one-click deployment with vLLM and autoscaling). This stack constitutes an alternative to Azure OpenAI Service or AWS Bedrock for enterprises primarily operating on Alibaba Cloud. [triangulated — Alibaba Cloud help docs + PAI-EAS blog + DashScope SDK doc]
International reach
International adoption is beginning to materialize. AI Singapore chose Qwen (over Llama) as the base for its regional model — a signal of adoption by a non-Chinese government actor. [triangulated — SiliconAngle + MIT Technology Review 2026 + Xinhua]
| Context | Recommendation | Why |
|---|---|---|
| Local deployment on smartphone or edge device | Qwen3-0.6B or 1.7B | The smallest open-weight range on the market, with reasonable performance for simple tasks. |
| Fine-tuning or domain adaptation on a single GPU | Qwen3-8B or 14B | Manageable on an A100 or two RTX 4090; base trained on 36T tokens, good transferability. |
| Enterprise production, performance/cost balance | Qwen3-30B-A3B (MoE) | 30B total, 3B active: performance close to the dense 32B, inference cost of a small model. |
| Intensive mathematical or code reasoning | Qwen3-235B-A22B or QwQ-32B | The flagship MoE for maximum capability; QwQ-32B for reasoning in a more modest footprint. |
| Cloud API integration, no own infrastructure | Qwen3-Max or Qwen3.5-Plus (DashScope) | Proprietary API-only, for use cases where managing infrastructure is not feasible. |
Strengths
Unmatched size coverage
Qwen3 covers 0.6B to 235B parameters in open-weight, with two architectures (dense/MoE). This is the most complete open range available in 2025: from edge (0.6B on a smartphone) to enterprise (235B MoE on a GPU cluster). Neither Llama 4 nor Mistral Large 3 covers this entire spectrum under an open license. [triangulated — qwenlm.github.io + arXiv 2505.09388 + datasciencedojo.com]
Derivative ecosystem dominance
As of March 2026, Qwen accumulates 942.1 million downloads on HuggingFace (versus 476 million for Llama). Qwen’s share of new HuggingFace fine-tunes grew from 1% in January 2024 to 69% in February 2026. Alibaba has more derivative models than Google and Meta combined. [triangulated — HuggingFace State of Open Source Spring 2026 + Xinhua + ATOM Report arXiv 2604.07190]
The 113,000 direct derivatives metric is already covered in the Ecosystem actors article — it is not developed further here.
Integrated reasoning without model switching
The integration of thinking and non-thinking modes into a single Qwen3 model is a notable architectural innovation. It simplifies deployment: a single infrastructure, a single API, two behaviors configurable via parameter. [replicated — qwenlm.github.io + Alibaba Cloud blog]
Massive multilingual pre-training corpus
36 trillion tokens for Qwen3 (versus 18T for Qwen2.5, 7T for Qwen2), across 119 languages — including synthetic data generated by Qwen2.5-Math and Qwen2.5-Coder for STEM domains. [triangulated — arXiv 2505.09388 + qwenlm.github.io]
Known limitations
Self-published benchmarks
Qwen’s performance comparisons (with DeepSeek-R1, Gemini-2.5-Pro, GPT-o1) are published by the Qwen team itself. The ATOM Report (arXiv 2604.07190) provides partial third-party validation, not systematic coverage. This limitation is shared with Llama (Meta) and Mistral — it is not a Qwen exception, but it deserves to be flagged. [replicated — arXiv 2604.07190 + BDTechTalks QwQ article]
Undisclosed training datasets
Qwen2.5 (arXiv 2412.15115) and Qwen3 (arXiv 2505.09388) technical reports describe token volumes and covered domains, but not the precise dataset composition: exact sources, synthetic-to-natural data ratio, deduplication or fairness metrics across 119 languages. Unlike DeepSeek, Alibaba has not published training cost figures. [triangulated — arXiv 2412.15115 + arXiv 2505.09388 + TechRxiv Qwen2.5 review]
Historical licensing heterogeneity (resolved in Qwen3)
Qwen1 and Qwen2 generations used the proprietary Tongyi Qianwen license — with a clause reserving commercial use for services below 100 million monthly active users, similar to the Llama 2 clause. Only models below 72B were Apache 2.0. This heterogeneity was resolved with Qwen3 (April 2025). [triangulated — GitHub Qwen issues #778 + HuggingFace Qwen2.5-72B discussions + CometAPI blog]
Geopolitics and export controls
Qwen models are trained in China in a context of US export controls on H100 chips and their successors. Alibaba responded, like DeepSeek, with algorithmic optimization rather than additional compute acquisition. The Apache 2.0 open-weight format complicates any post-distribution access restriction: once a model is published, its global usage cannot be controlled. [triangulated — MIT Technology Review 2026 + Congress.gov CRS R48642 + AI Frontiers]
Recent organizational instability
The March 2026 restructuring (creation of “Alibaba Token Hub”) and leadership changes signal a period of transition. The impact on development cadence and open-source commitment is not yet assessable. [single source — Wikipedia Qwen, to be confirmed]
Key takeaways
In three years, Qwen moved from a product captive within the Alibaba ecosystem to the world’s most-derived open-weight model family. Four elements explain this trajectory:
- The range: no other actor covers open-weight 0.6B to 235B with two architectures under Apache 2.0.
- The cadence: major releases every six to nine months, with arXiv technical reports at each generation.
- The Apache 2.0 strategy: by fully opening the license on open-weight models since Qwen3, Alibaba removed the primary friction to commercial adoption.
- Vertical integration: ModelScope + DashScope + PAI-EAS creates a complete chain from raw weights to managed cloud deployment, mirroring what AWS or Azure offer with their respective partner models.
What remains uncertain: the sustainability of the open-weight commitment under commercial pressures, alignment quality across 119 languages (unpublished), and the consequences of the 2026 organizational restructuring.