In brief

Hugging Face is often reduced to its model-sharing platform. That is reductive. The Franco-American company founded in 2016 operates on three distinct levels: Hugging Face the company (400+ employees, revenues ~$130M in 2024), the Hub (the platform for depositing and distributing models, datasets and applications), and open-source libraries (transformers, datasets, diffusers, accelerate and six others — usable without an account, freely distributed on PyPI). These three levels reinforce each other but are not the same thing: you can use transformers without ever creating a Hub account, and you can host a model on the Hub without ever touching the libraries.

What distinguishes Hugging Face from all other AI ecosystem players: its multi-provider neutrality. Llama (Meta), Gemma (Google), Mistral, Qwen (Alibaba), DeepSeek — all distribute their models via the Hub. This role as neutral infrastructure, accepted by direct competitors, is structurally difficult to replicate.


The origin: from a teen chatbot to global infrastructure

To understand Hugging Face, you need to start with its origin — which has nothing to do with what the company has become.

In 2016, three Frenchmen — Clément Delangue (CEO), Julien Chaumond (CTO) and Thomas Wolf (CSO) — met through an online Stanford course and launched a chatbot application aimed at teenagers. The 🤗 emoji that gives the company its name reflected the ambition of the time: a “virtual friend,” accessible and playful, deliberately going against corporate seriousness. Technically functional, the product never found its market.

The turning point came in late 2018, with Google’s publication of BERT. The Hugging Face team produced a PyTorch implementation in one week and published it as open source. The community reaction was immediate: thousands of practitioners who wanted to experiment with BERT but lacked resources to reproduce the original code rushed to the implementation. Julien Chaumond, in a Contrary Research interview, summarized: “That moment clarified our direction.”

That direction: becoming the rapid distribution relay for research advances to practitioners. Not a lab producing models — infrastructure making them accessible.

In 2019, after a stint in an accelerator (Betaworks), the team officially abandoned the chatbot. The transformers library was launched under the tagline “GitHub for ML.” In 2020, the Hub opened as a central model registry — the idea being to treat model weights like Git repositories: versioned, shareable, downloadable in a single line of code.

This cultural imprint — accessibility, zero friction, against corporate secrecy — persists in the platform’s DNA today.


The fundamental distinction: company, Hub, libraries

Before going further, a clarification that conditions everything else. When someone says “I use Hugging Face,” they may mean three very different things:

Hugging Face the company: a commercial entity with 400+ employees headquartered in Brooklyn, funded to the tune of $395M cumulatively (including $235M in August 2023 in a Series D involving Google, Amazon, Nvidia, Intel, AMD, Qualcomm, IBM and Salesforce). It generates revenues through Enterprise services and cloud inference.

The Hub: the web platform accessible at huggingface.co. It is a registry of Git repositories specialized for ML artifacts — models, datasets, interactive applications. You create an account, upload files, browse millions of public contributions. The exact metaphor: GitHub, but for neural network weights and datasets.

The libraries: Python code distributed via PyPI. transformers, datasets, diffusers, accelerate, peft, and others. Usable without connectivity, without an account, without dependency on the Hub. A developer can integrate from transformers import AutoModelForCausalLM into their code and never open huggingface.co.

These three levels feed each other: the libraries create stickiness independent of the platform, the Hub facilitates sharing that drives library usage, and the company monetizes professional uses of both. But they are technically separable — and this separability is precisely what makes the business model both powerful and fragile.

In plain terms: Hugging Face is like if GitHub (the repository platform), npm (the package manager) and a cloud services company were all one. Each layer generates value for the others — and you can use one without the other two.


The Hub: anatomy of a distribution platform

The Hub is the centerpiece of the visible offering. Its architecture organizes around three types of repositories.

Models

The central repository holds 2M+ public models as of spring 2026 — an impressive figure that demands an immediate caveat: concentration is extreme. According to the official HF “State of Open Source on Hugging Face: Spring 2026” report, a tiny fraction of models captures the bulk of downloads. This concentration is not an anomaly — it is the classic power law of distribution platforms.

What matters more than raw volume is coverage: Llama (Meta), Gemma (Google), Mistral, Qwen (Alibaba), DeepSeek, Falcon — all major open-weights models transit through the Hub. For a developer wanting to test a recently announced model, the Hub is reflex number one. Qwen counts 113,000 derivative models on the Hub — more than Meta and Google combined — a sign that the community fine-tunes and redistributes massively from base weights.

Each model comes with a Model Card — a standardized document describing intended use, known biases, training data (when documented), and limitations. This format, initiated by Hugging Face, has become the de facto convention in the community. The EU AI Act cites it as a documentation best practice.

Datasets

500,000+ public datasets, with collaborative labeling tools integrated since the acquisition of Argilla in June 2024 (~$10M). The strategic signal of this acquisition is clear: HF is betting that data quality will be the next differentiation vector — beyond simply distributing pre-trained models. Controlling data curation means controlling the quality of models that emerge from it.

Spaces

500,000+ interactive applications deployed on the Hub, powered by Gradio (acquired in late 2021) or Streamlit. Spaces allow anyone to deploy a model demonstration accessible via a simple browser — without a backend, without ops. For a researcher publishing a paper, it is the fastest way to make a result interactive and reproducible.

Free Spaces run on shared CPU. For GPU-intensive uses, a pay-as-you-go “dedicated hardware” system allows upgrading for a few euros per hour. This is one of the rare direct monetization points of the Hub for individual users.

In plain terms: the Hub works like a vast global bazaar of AI — millions of public contributions, a standardization effort (Model Cards), and an interactive showcase layer (Spaces) that makes models testable without installation.


The libraries: the technical ecosystem that creates loyalty

If the Hub is the storefront, the libraries are the cement. Their adoption creates technical dependency that the platform alone would not provide.

LibraryRoleAdoption signal
transformersAccess to pre-trained models (NLP, vision, audio, multimodal)160,000 GitHub stars (April 2026), 300,000 PyPI downloads/day, 1M+ checkpoints on Hub
datasetsLoading and processing ML datasetsDe facto standard in academic research
diffusersImage/video/audio generative modelsDominant for diffusion models
accelerateMulti-GPU/TPU/distributed trainingTraining abstraction layer
peftEfficient fine-tuning (LoRA, QLoRA, Prefix Tuning)Open-source post-training standard
tokenizersHigh-performance tokenization (Rust)Integrated in transformers
safetensorsSecure weight loading format (replaces pickle)Growing adoption vs .pt format
trlRLHF/DPO/PPO fine-tuningRLHF post-training standard
evaluateStandardized evaluation metricsUsed in Open LLM Leaderboard

Transformers v5: a major evolution

In March 2026, Hugging Face published version 5 of transformers — the first major version in five years. 1,200 commits since the last minor release. The central architectural novelty: native support for hybrid architectures combining classical attention and Mamba-type layers (SSM), reflecting the diversification of LLM architectures in 2025-2026. The release cadence also changed: one minor version per week announced, compared to one per five weeks previously.

The promise was kept. Six months later, in early September 2026, branch 5 is on its sixteenth minor version — exactly the announced rhythm. This is worth noting, because an announced release cadence is rarely held: it usually runs into the cost of backward compatibility. The content of those releases also says something about the library’s place in the ecosystem — support for Muse Glimmer lands in the version published shortly after Meta’s announcement.

The popularity of transformers is not incidental to HF’s health: it is the library that anchors developers in the HF ecosystem. A developer who has structured their code around from transformers import ... does not easily migrate to an alternative — even if the Hub became paid or disappeared.

In plain terms: HF libraries are to AI what jQuery was to the web in 2010 — an abstraction layer so widely adopted that it shapes how developers think about the problem, not just how they implement it. transformers counts 160,000 GitHub stars and 300,000 downloads per day.


Cloud services: from free to Enterprise

Serverless inference (free)

The Inference API allows querying thousands of models via HTTP request, without own infrastructure. Ideal for prototyping and low-traffic applications. The limitation: very popular models are subject to quotas and queues. This free tier is a showcase — it convinces developers the Hub works, before converting them to paid offerings.

Inference Endpoints (paid, dedicated)

Secure deployment with auto-scaling (including scale-to-zero to pay only during use), choice of cloud (AWS, Azure, GCP), region, and hardware. PrivateLink support for companies that do not want to route their data through the public internet. Base rate: starting from $0.033 per hour on CPU, significantly more on GPU.

Enterprise Hub

This is the core of HF’s business model. The Enterprise Hub (formerly Private Hub) is a private version of the Hub with:

  • SSO and RBAC (role-based access control) for large teams
  • Audit logs for compliance
  • Regionalized storage for GDPR compliance and data sovereignty requirements
  • Data governance to control who accesses what

2,000+ client organizations, including Intel, Pfizer, Bloomberg, eBay. 30%+ of the Fortune 500 uses the Hub according to HF communications. Pricing starting at $20 per user per month for small teams.

The 2,000 organizations figure comes from official HF communications (pricing page). Exact revenues are not published, but third-party sources (Sacra, Latka) estimate 2024 ARR between $130M and $160M — with a significant portion coming from Enterprise services.


Open-R1 and the research dimension

Hugging Face is not just a distribution platform — it also conducts reproducible research. The most striking example is Open-R1, launched in January 2025 by Leandro von Werra, HF head of research.

Following DeepSeek R1’s publication under MIT license, HF launched a project for entirely transparent replication: reproducing not just the weights, but also the training data and complete pipeline, to make “pure reinforcement learning” reasoning accessible and verifiable by the community. 10,000 GitHub stars in three days — a signal of strong community appetite for reproducibility.

The resource mobilized: the HF Science Cluster, consisting of 768 H100 GPUs. This is a significant asset — this level of compute allows HF to conduct reproducible frontier experiments that academic teams without dedicated cloud cannot perform. Similar projects followed: SmolLM (series of efficient small models), FineWeb (massive open-source pre-training dataset).

This research dimension plays an important role in HF’s credibility: the company does not simply distribute — it contributes to the shared knowledge base.

In plain terms: Open-R1 is the concrete example of what “neutral but active” means for HF. Rather than taking sides between labs, HF makes reproducible what labs have accomplished — by surfacing training data and code to the level where anyone can inspect them.


Kernel Hub: extending toward compute

In 2025, HF launched the Kernel Hub — a repository of optimized GPU kernels (NVIDIA, AMD). The logic is a natural extension toward the compute layer: after standardizing the sharing of models and datasets, standardize the sharing of low-level optimizations.

The Kernel Hub accepts kernels written in CUDA, Triton or other GPU programming languages. For developers who spend time optimizing matrix operations or attention kernels, it is a community reuse library. For HF, it is a foothold in hardware optimization — a domain until now dominated by chip manufacturers.

It is too recent to have significant adoption data. But the direction is clear: HF is moving up toward the infrastructure layer, not just the model layer.


Community as infrastructure

The Hub is not just a repository — it is a social space. Several community products are integral to it.

The Open LLM Leaderboard, developed with EleutherAI, has become the de facto reference for comparing open-source models on reproducible benchmarks. Version 2, launched in March 2025, integrates more robust benchmarks: IFEval, GPQA, MATH, BBH, MMLU-Pro. 2 million unique visitors over 10 months, 300,000 active members monthly. Whoever controls the reference evaluation influences quality perceptions across an entire community.

Daily Papers: daily curation of arXiv papers by the community. Hundreds of papers submitted and evaluated each day. For tens of thousands of practitioners, this is the daily filter for ML research. This curation role is not directly monetized — it is an engagement and traffic asset.

HuggingChat: an open-source alternative to ChatGPT, allowing interaction with models hosted on the Hub directly from a browser. Modest usage compared to large proprietary models, but useful as a showcase for open-weights models.


Structural strengths

Dominant network effect. The Hub benefits from the classic network effect: more models attract more users, who deposit more models. With the critical mass reached, switching costs are real — a developer who has their models on HF, their pipelines built with transformers, their evaluations on the Open LLM Leaderboard does not easily migrate.

Multi-provider neutrality. Neither OpenAI, Google, nor Meta can host their competitors’ models. HF does, and this is accepted by all because it is structurally in everyone’s interest to have neutral distribution infrastructure. This neutrality is difficult to replicate by a player developing its own models — the conflict of interest is obvious.

Libraries as independent stickiness. transformers creates dependency on the HF ecosystem that persists even if someone does not use the Hub. This is a rare form of double “moat”: the platform AND the tooling.

European political dimension. HF is the most visible actor associated with “French Tech AI” — its co-founders are French, BLOOM was trained on the Jean Zay supercomputer (GENCI). In discussions on European digital sovereignty, HF is perceived as “neutral” infrastructure that can host models without dependency on an American or Chinese hyperscaler. This dimension is not trivial for institutional partnerships.


Structural tensions

The free infrastructure business model. HF publishes its core libraries for free, hosts public models for free, and absorbs bandwidth costs for the most popular ones. Download concentration creates an asymmetry: heavily used models (often from large organizations that monetize these models in their products) generate infrastructure costs for HF without direct contribution. This freemium model works as long as Enterprise revenues offset public hosting costs. If volumes continue growing faster than revenues, the tension intensifies.

Progressive hyperscaler disintermediation. Vertex AI Model Garden (Google) and Azure AI Studio (Microsoft) integrate HF models directly. This appears favorable — but it is also an alternative: a developer can access HF models via Azure without ever opening huggingface.co. If hyperscalers make integration sufficiently transparent, the Hub becomes optional in a professional workflow. The threat is not frontal — it is slow disintermediation.

ML supply chain security. The Hub’s attack surface is significant: Palo Alto Networks Unit 42 documented “namespace reuse” attacks — malicious models published under names similar to popular models. HF has deployed automatic security scanners, but publication speed (thousands of models per day) structurally exceeds manual verification capacity. ML supply chain security is a recurring criticism vector, likely to intensify as HF models integrate into critical production pipelines.

Unilateral governance. The Hub is private infrastructure presented as neutral. HF decides moderation rules, banned models, hosting conditions. This asymmetry between the promise of neutrality and effective control is a tension rarely discussed publicly, but structurally significant for organizations building critical dependencies on the platform.


Decision matrix: when to use what in the HF ecosystem

NeedHF componentWhy
Load a pre-trained model in Python codetransformers (PyPI)One line of code, 1M+ checkpoints available, no account needed
Test a model interactively without installationHub Spaces or HuggingChatBrowser demonstration, zero technical friction
Share a model or dataset with the communityHub (public repository)Git versioning, integrated Model Card, indexed discoverability
Deploy a model in production with SLAInference Endpoints (paid)Auto-scaling, PrivateLink, cloud/region choice
Private team with access control and complianceEnterprise HubSSO, RBAC, audit logs, regionalized storage
Fine-tune with LoRA/QLoRA without custom codepeft + trlOpen-source standards, abundant examples
Generate images/videodiffusersDominant architecture for diffusion models

Some practical heuristics

  • Start with transformers, not the Hub. The library is independent and sufficient for 80% of experimentation use cases.
  • Use the Hub for distribution, not for production inference. The free Inference API is for prototyping. In production, use Inference Endpoints or your own infrastructure.
  • Model Cards: read them before deploying. They contain warnings about biases and discouraged uses — information that proprietary lab Model Cards do not always publish.
  • Enterprise Hub for organizations with GDPR constraints: storage regionalization covers data localization obligations.

What we retain in 2026

Hugging Face has become something quite rare in the AI ecosystem: infrastructure accepted by all because neutral toward all. Google, Meta, Alibaba and dozens of competing startups distribute their models through the same Hub — not out of naivety, but because the alternative (building your own community distribution platform) costs more than contributing to shared infrastructure.

The strength of this position is real. So is its fragility: the free infrastructure economic model requires Enterprise revenues to grow faster than hosting costs, hyperscalers to remain integrators rather than substitutes, and community trust in governance neutrality to hold.

transformers v5 at 160,000 GitHub stars, Open-R1 reproducible by the community, Kernel Hub for low-level optimizations: HF does not merely distribute — it increasingly anchors the technical layer on which open-source AI runs.