In brief

Nvidia holds approximately 92% of the AI accelerator market for training in 2025. This dominance does not rest solely on the raw performance of its chips — it also relies on twenty years of software ecosystem built around CUDA, a proprietary programming model that competitors have not yet managed to work around seriously. But an erosion is underway, through an unexpected vector: inference and the custom chips of major cloud providers.


Why a GPU for AI?

A standard processor (CPU) is designed to execute complex instructions one after the other, with at most a few dozen cores. A GPU works the opposite way: it packs thousands of simple cores capable of processing many operations in parallel.

Training or running a large language model essentially amounts to multiplying matrices — arrays of numbers — millions of times per second. GPUs are naturally suited to this task. Nvidia’s H100 integrates 16,896 compute cores. The B200, released in 2024, groups the equivalent via a two-die architecture linked together.

These cards also embed very high-speed memory, HBM (High Bandwidth Memory), stacked directly on the chip. The H100 has 80 GB at 3.35 TB/s of bandwidth. The B200 goes up to 192 GB at 8 TB/s. This fast memory is the dimensioning factor for large models: the heavier a model, the more memory is needed to run it without degrading performance.


CUDA: the real barrier to entry

In 2006, Nvidia launched CUDA (Compute Unified Device Architecture), a programming model that allows developers to write C/C++ code executed directly on Nvidia GPUs. This is not a simple interface — it is an entire ecosystem.

Around CUDA orbit highly optimized libraries: cuBLAS for linear algebra, cuDNN for deep learning primitives, NCCL for GPU-to-GPU communication within a cluster. These libraries are co-developed with the hardware at each generation. When a new chip launches, CUDA libraries exploit it immediately. No competitor can offer this level of software-hardware integration.

The result: over 4 million developers trained on CUDA, 3,000 GPU-accelerated applications, and all major frameworks — PyTorch, TensorFlow, JAX — relying on CUDA under the hood. An engineer writing PyTorch is using CUDA without knowing it.

Migrating to an alternative (AMD’s ROCm, for instance) represents several months of engineering per organization, with no guarantee of equivalent performance. ROCm remains 10 to 30% less performant than CUDA on compute-intensive workloads, according to 2025 benchmarks.

In short: Nvidia chips can be matched within hours. The twenty years of optimized libraries, 4 million trained developers, and silent integration into PyTorch cannot be matched. The real lock-in is not the silicon, it is the software stack.


Hardware generations: Ampere, Hopper, Blackwell

A100 (Ampere, 2020)

The A100 is the chip that powered the first wave of large-scale LLM deployments. 80 GB of HBM2e, 2 TB/s of bandwidth, NVLink 3.0 interconnect at 600 GB/s between GPUs. It introduces third-generation Tensor Cores, specialized units for matrix multiplications at reduced precision (FP16, BF16), delivering a compute-per-watt ratio well above standard cores.

H100 (Hopper, 2022)

The H100 marks a qualitative leap. Manufactured at 4 nm at TSMC, it integrates 80 billion transistors and introduces the Transformer Engine — a dedicated unit for FP8 (8-bit floating point) computation, allowing inference throughput to be doubled compared to BF16. In training, gains measured by MLPerf place the H100 at 2 to 3 times the performance of the A100. In LLM inference, the factor reaches 10–20×, depending on the precision used. A single H100 unit price is around $37,000 on the market.

NVLink 4.0 interconnect reaches 900 GB/s, and NVL configurations (up to 8 GPUs on a DGX H100) allow running models of several hundred billion parameters while keeping the whole model in shared memory.

B200 (Blackwell, 2024)

The B200 crosses a physical limit: the maximum surface area of a lithography reticle (approximately 850 mm²) was no longer sufficient to integrate all the transistors desired on a single die. Nvidia adopts a multi-die architecture — two dies connected by an internal bus at 10 TB/s. The result: 208 billion transistors, 192 GB of HBM3e at 8 TB/s, NVLink 5.0 at 1,800 GB/s.

The announced gains — ×5 in FP4 compute versus the H100 — rely on FP4 precision (4 bits), implying specific quantization techniques. In training, the measured gain on DGX systems is around ×3. In inference, ×15.

The scaling comes at a cost: NVL72 racks (72 B200 GPUs interconnected) consume over 120 kW. They require liquid cooling and special electrical infrastructure. Several Blackwell deployments were delayed not by chip shortages, but by power supply constraints in datacenters.

In short: each generation doubles to triples raw performance, but power consumption climbs at the same pace. A Blackwell NVL72 rack consumes as much as a residential neighborhood — hence the pressure on datacenters and the mandatory liquid cooling.


GenerationYearHBM memoryBandwidthTypical drawTraining gain vs previousInference gain vs previous
A100 (Ampere)202080 GB HBM2e2 TB/s~ 400 Wbaselinebaseline
H100 (Hopper)202280 GB HBM33.35 TB/s~ 700 W× 2 to 3× 10 to 20 (FP8)
B200 (Blackwell)2024192 GB HBM3e8 TB/s~ 1,000 W× 3× 15 (FP4)

In plain terms: each generation doubles or triples raw performance, but power draw climbs at the same rate. A Blackwell NVL72 rack consumes as much as a residential block — hence the pressure on datacenters, and mandatory liquid cooling.


The challengers: AMD, Intel, and in-house chips

AMD is the only serious competitor on the merchant GPU market, with 5 to 8% market share. The heaviest competition is not sold there: the hyperscalers’ in-house chips — Google’s TPU, Amazon’s Trainium — serve their own infrastructure and never reach that market. Its MI300 lineup (MI300X: 192 GB of HBM3) is 20 to 30% cheaper than the H100, which has earned it deployments at Meta, OpenAI, and on supercomputers like El Capitan. But the ROCm software lag remains the primary barrier.

Intel is below 1% of the market. Its Gaudi 3 accelerator targets low-cost inference, without establishing itself as a credible alternative at scale.

Hyperscaler in-house chips are the real structural threat. Google uses its TPUs (v7 “Ironwood”) to train and run its Gemini models, claiming cost reductions of 40 to 60% compared to external GPUs. AWS has Trainium 2 (used to train certain customer models). Microsoft is developing Maia 2, Meta MTIA. These chips are not sold — they capture the internal demand of hyperscalers. Projections put them at 15 to 25% of the total market by 2026–2027, primarily in inference.


The real debate: erosion through inference

Training new models remains the domain where Nvidia is hardest to replace: it requires great flexibility, high precision (BF16, FP32), and complex management of inter-GPU communications. Nvidia holds approximately 90% of this segment.

Inference — running an already-trained model to respond to user queries — is structurally different. A model in production has a fixed topology. It can be optimized, quantized, specialized. This is where custom ASICs are economically justifiable despite their high development costs.

In short: Nvidia is not losing the war, but the ground is shifting. Training remains its stronghold (90%). Inference is being attacked — not by AMD/Intel on the open market, but by hyperscalers who are internalizing (Google TPU, AWS Trainium). The CUDA lock-in is loosening slowly from below, via Triton and JAX.

If TPUs and Trainium capture hyperscaler inference, they do not directly compete with Nvidia on the open market — they simply reduce the addressable demand. The impact on Nvidia’s revenue depends on total market growth: if AI usage grows faster than hyperscaler internalization, Nvidia can lose market share and gain in absolute value.

Another erosion vector is more discreet: Triton, an open-source language developed by OpenAI allowing portable GPU kernels to be written, compilable on AMD GPUs and Google TPUs. New kernels written in Triton are not tied to CUDA. As codebases renew, the stock of exclusively CUDA code diminishes.


Key takeaways

  • Nvidia GPUs dominate AI compute at ~92% in training, through a combination of performant hardware and twenty years of CUDA software ecosystem.
  • Each generation brings measurable gains: H100 ×2–3 vs A100 in training, ×10–20 in inference; B200 ×3 in training, ×15 in inference vs H100.
  • Performance progression is matched by an equally rapid progression in power consumption: Blackwell racks exceed 120 kW, making liquid cooling mandatory.
  • The erosion of Nvidia’s dominance comes primarily from hyperscaler in-house chips (TPU, Trainium) on inference — not from AMD.
  • The “CUDA lock-in” is real but not permanent: Triton and JAX open portability paths for new projects.