In brief
Nvidia controls between 80 and 92% of the AI accelerator market depending on the metric used. But since 2023, several players have managed to deploy alternatives in real production — not just on paper. AMD runs at Meta and Microsoft, AWS builds its own chips to train Claude models, Cerebras signed a ten-billion-dollar deal with OpenAI. The problem is less raw chip performance than the software ecosystem surrounding them — and that is where Nvidia’s dominance is hardest to dislodge.
Three families of alternatives
Before comparing players, three radically different types of approaches must be distinguished.
Open-market GPUs (AMD Instinct, Intel Gaudi) are chips that any company can purchase. They compete directly with Nvidia GPUs on the same use cases — training and inference of large models — but can run on any cloud or server. This is the most contested market.
Proprietary ASICs (AWS Trainium, Google TPU, Microsoft Maia) are custom-designed circuits built by large cloud platforms for their own needs. They are not available for purchase: they power internal services and, in some cases, the cloud offerings of those platforms. Their advantage: cost and energy efficiency optimized for specific workloads. Their limitation: no flexibility outside the intended use case.
Alternative architectures (Cerebras, Groq, SambaNova, Tenstorrent) invert the problem. Rather than improving the classic GPU, they propose radically different designs: a processor covering an entire silicon wafer, a language processing unit with deterministic scheduling, reconfigurable compute units. Significant gains on targeted workloads, significant rigidity on others.
AMD: the only credible challenger on the open market
In 2023, AMD was marginal in production AI. By 2026, Meta runs 100% of its Llama 405B model traffic on AMD MI300X chips. Microsoft Azure and Oracle Cloud offer MI300X instances in general availability.
This shift stems from an architectural decision. The MI300X carries 192 GB of HBM3 memory — versus 80 GB for Nvidia’s H100 — with 5.3 TB/s bandwidth. For very long-context models, which must retain large amounts of data in memory during inference, this advantage is decisive.
The MI350 generation (2025) pushes further: 288 GB of memory, 8 TB/s bandwidth. MLPerf Training v5.1 benchmarks show performance close to Nvidia in FP8 precision, with gains up to 2.8x over the previous generation.
AMD’s problem is not the silicon. It is ROCm, its equivalent of CUDA. Despite ROCm version 7 (September 2025), practical measurements show CUDA outperforms ROCm by 10 to 30% on compute-intensive workloads, even where AMD holds a theoretical advantage on paper. Porting a CUDA application to ROCm remains significantly more complex than AMD publicly communicates.
In short: AMD has the memory (192 → 288 GB versus 80 GB for the H100), AMD has the price (-20 to -30%), AMD has serious customers (Meta, Microsoft). But AMD does not have twenty years of accumulated software optimization. For a new greenfield project, ROCm may suffice. To migrate an existing CUDA codebase, the engineering cost remains prohibitive.
Captive ASICs: when hyperscalers build their own chips
The clearest strategy is at AWS. Project Rainier is a dedicated facility in Indiana: 500,000 Trainium2 chips, on a 1,200-acre site, devoted exclusively to training Anthropic’s Claude models — with five times the compute power of previous generations.
The Trainium3 chip, launched in late 2025 using TSMC’s 3nm process, delivers 2.52 petaflops FP8 per chip, 144 GB of HBM3e, and 4.9 TB/s bandwidth — 4.4 times more compute capacity than Trainium2. AWS quotes Trn2 instance costs approximately two times lower than comparable H100 instances for adapted workloads.
This model raises a structural question: captive ASICs do not merely compete with Nvidia. They also divert volumes that might otherwise have gone to AMD or other challengers. Hyperscalers are self-supplying, shrinking the addressable market for all GPU manufacturers — including Nvidia’s alternatives.
Google has followed the same logic with its TPUs since 2016, with the distinction that TensorFlow then JAX were developed primarily for TPU. Meta experiments with its own inference chips (MTIA). Microsoft launched the Maia 100, but its production deployment remains limited to internal use — and the next generation (Maia 200) has suffered significant delays, pushed to 2026 after design changes and integration difficulties.
In short: a captive ASIC does not attack Nvidia head-on — it subtracts volumes from the addressable market. When Anthropic trains Claude on AWS Trainium2, that is as many H100 chips that AWS does not buy from Nvidia. Nvidia’s dominance erodes silently from below, without a frontal battle.
Alternative architectures: niche performance, real risks
Cerebras Systems has built something unusual: the Wafer-Scale Engine (WSE-3) is a chip covering an entire 300mm silicon wafer — where a classic GPU occupies only a fraction of one. The result: 4 trillion transistors, 900,000 AI cores, and 44 GB of SRAM embedded directly on the chip. The advantage: data does not need to transit through external memory (HBM), the classic GPU bottleneck. In January 2026, OpenAI signed with Cerebras a partnership estimated at more than $10 billion for 750 MW of compute capacity through 2028.
Groq had developed an LPU (Language Processing Unit) with deterministic scheduling — unlike GPUs, which handle tasks in parallel in a probabilistic manner. On Llama 70B inference, the LPU reached 276 to 300 tokens per second in standard configuration, up to 1,665 tokens per second with acceleration techniques. The limitation: exclusive optimization for inference, which reduced the addressable market. Nvidia acquired Groq in late 2025 for approximately $20 billion.
Tenstorrent, led by Jim Keller (chip architect at AMD, Apple, and Intel), bets on a radically different strategy: a fully open-source software stack, RISC-V architecture. The Blackhole chip (6nm TSMC) lines up 140 specialized cores, 774 TFLOPS in FP8. With $700 million raised and participation from Jeff Bezos, the company has assembled a 192-chip production cluster.
SambaNova offers RDUs (Reconfigurable Data Units) on its SN50 system (2026) and claims speeds five times faster than competing GPUs and a total cost three times lower for agentic AI workloads and MoE-type models such as DeepSeek-R1. These figures come from company communications and have not been verified by independent benchmarks.
The real battle: CUDA, not silicon
Jim Keller described CUDA as a “swamp, not a moat.” The argument: like x86 at Intel in the 2000s, technical legacy eventually becomes a burden rather than an advantage.
The opposing thesis is defended by the majority of financial analysts. CUDA represents twenty years of cumulative investment: more than 4 million trained developers, more than 3,000 optimized applications, and libraries such as cuDNN and TensorRT deeply integrated into production pipelines. CUDA migration tools to other environments exist but are considered insufficient for production deployments.
In practice, the performance gap between CUDA and alternative software is real and documented. It is not a hardware gap — AMD has, on certain metrics, a 32% theoretical advantage over Nvidia. It is a gap of software optimization accumulated over years.
The hyperscaler strategy to bypass this problem is to go through PyTorch, which has become the industry standard. Google and Meta push native PyTorch backends for TPU and AMD (TorchTPU, ROCm via PyTorch) — hoping the abstraction layer will make the underlying silicon progressively interchangeable.
In short: PyTorch is the discreet strategic weapon of challengers. If each chip has an optimized PyTorch backend, the user no longer touches CUDA directly — the dependency disappears at the API level. This is the same strategy as Java versus x86 in the 2000s: abstract the underlying hardware.
Dashboard — where things stand
| Player | Type | Real deployments | 2025–2026 signal |
|---|---|---|---|
| AMD Instinct MI300X/MI350 | Open-market GPU | Meta (Llama 405B), Microsoft Azure, Oracle Cloud | Primary challenger, ROCm software lag |
| AWS Trainium2/3 | Captive ASIC | Claude training (Anthropic), Trn2 instances | Project Rainier 500K chips, Trainium3 at 3nm |
| Google TPU | Captive ASIC | Internal Google services, Google Cloud | Since 2016, native TensorFlow/JAX |
| Microsoft Maia | Captive ASIC | Internal use only | Maia 200 delayed to 2026 |
| Intel Gaudi 3 | Open-market GPU | IBM Cloud, Dell PowerEdge | Line discontinued — adoption risk |
| Cerebras WSE-3 | Wafer-scale | OpenAI partnership | $10B deal signed January 2026 |
| Groq LPU | LPU inference | Public Groq cloud | Acquired by Nvidia late 2025 (~$20B) |
| Tenstorrent Blackhole | RISC-V + Tensix | 192-chip cluster | Open-source, $700M raised |
| SambaNova RDU | Reconfigurable | Niche MoE/agentic | Figures not independently verified |
| Huawei Ascend | Closed-market GPU | Dominant in China | Outside Western market (export restrictions) |
Key takeaways
- Nvidia maintains 80 to 92% of the market depending on the metric. In units deployed for inference, its share would be 60 to 75% including hyperscaler ASICs.
- AMD is the only challenger with real large-scale production deployments (Meta, Microsoft, Oracle). Its memory advantage on long-context models is documented. Its software lag (ROCm vs. CUDA) remains significant.
- Captive hyperscaler ASICs (AWS Trainium, Google TPU) represent an alternative that bypasses all open-market players — including Nvidia’s challengers.
- Intel has abandoned its Gaudi line, leaving a gap in the open-market GPU segment and creating real risk for companies that had begun adopting it.
- Nvidia’s dominance rests less on silicon than on the CUDA software ecosystem — twenty years of investment that no challenger has yet matched in production conditions.
- The market is growing faster than challengers are gaining share: a decline in Nvidia’s market share percentage can correspond to an increase in its revenue in absolute terms.