In brief

Since 2022, five startups — Cerebras, Groq, SambaNova, Etched, and Tenstorrent — have designed specialized chips for large language model (LLM) inference. Their shared bet: where Nvidia GPUs are optimized for training (billions of parameters computed in parallel over large batches), inference follows a radically different logic — sequential, latency-sensitive, constrained by memory bandwidth. The picture in 2026: Cerebras claims up to 2,500 tokens per second per user on Llama 3.1 405B, independent data confirms Groq at 877 tokens/s on Llama 3 8B. But the picture is more nuanced: Groq was acquired by Nvidia in December 2025 and is no longer an independent challenger. Etched has raised $620 million without delivering a single chip to an external customer. And Tenstorrent reaches only 50% of its theoretical peak in real LLM inference.

This article covers the five inference ASIC startup players. For hyperscalers (AMD, AWS Trainium, Google TPU, Microsoft Maia), see the alternative chips article.


Understanding the problem: why the GPU isn’t the right chip for inference

Picture a workshop that manufactures furniture. For an order of 1,000 identical tables, a large shop with 200 workers operating in parallel is unbeatable — that’s LLM training. But if a customer walks in and asks for a single custom chair while waiting at the counter, those 200 workers don’t help: the machine is too large for a sequential task, and the customer still waits for the first worker to finish the first step before starting the next. That’s LLM inference.

The memory bandwidth bottleneck

LLM inference generates one token at a time, autoregressively. At each step, the model must read all of its parameters from memory, plus a context cache (the KV-cache, Key-Value cache) that grows with each generated token. On an Nvidia H100 GPU, HBM memory bandwidth is 3.35 TB/s. For a Llama 3.1 70B model in FP16, that’s approximately 140 GB of weights to read at each step — a theoretical minimum of about 42 ms per token, meaning roughly 24 tokens per second for a single user in decode mode. [source: IEEE Spectrum 2024 “Challengers Are Coming for Nvidia’s Crown”, arxiv 2503.11698]

The structural problem: GPUs were designed for training, which maximizes compute unit utilization (matmul over large batches). Autoregressive inference cannot parallelize over future tokens — they haven’t been generated yet. GPUs massively underutilize their compute capacity at low-batch inference.

Three distinct metrics — often confused

Vocabulary

  • Throughput (aggregate tokens/s): total tokens per second across all requests combined. Datacenter and server cost metric. Improves with batching — the more requests are grouped, the more efficiently the chip is used.
  • Per-user latency (tokens/s per user): throughput as seen by a single interactive user. Does NOT improve with batching — on the contrary, more parallel requests increase per-user latency. This is the conversational metric.
  • Time to first token (TTFT): delay before the first token appears. Critical for interactive applications.
  • Cost per token: infrastructure cost per produced token. Depends on hardware total cost of ownership (TCO) and energy consumption.

In plain terms: when a vendor announces “1,000 tokens/s,” the decisive question is: is this 1,000 aggregate tokens across 50 simultaneous requests (= 20 tokens/s per user) or 1,000 tokens per user in direct interaction? The difference is 50×. Both metrics have their place, but they are not interchangeable.


Cerebras WSE-3 — the entire silicon wafer

The founding idea: eliminate the external memory bus

Cerebras chose the most radical solution to the bandwidth bottleneck: eliminating external memory entirely. The WSE-3 (Wafer-Scale Engine 3, 2024) occupies the full surface of a 300mm silicon wafer — 46,225 mm², versus 814 mm² for an H100 (57 times larger). It integrates directly:

  • 4 trillion transistors (5nm TSMC process) [source: Cerebras press release March 2024, confirmed IEEE Spectrum]
  • 900,000 AI cores (Sparse Linear Algebra cores) [source: Cerebras.ai/chip, Hot Chips 2024]
  • 44 GB of on-chip SRAM — memory lives on the chip itself [source: arxiv 2503.11698]
  • Memory bandwidth: 21 petabytes/s — approximately 7,000 times the HBM bandwidth of an H100 [source: Cerebras datasheet]

The key idea: by eliminating the external memory bus, the WSE-3 removes the bottleneck that limits GPUs in inference. Where an H100 must fetch weights from HBM at 3.35 TB/s, the WSE-3 reads them from on-chip SRAM at 21 petabytes/s — a difference of roughly 3 to 4 orders of magnitude.

In plain terms: an H100’s HBM memory is like a library in an adjacent room — you have to leave your desk to fetch books. Cerebras’s on-chip SRAM is the library at your desk, on the same shelf, within reach. The access speed difference is astronomical.

Benchmarks and triangulation

ModelTokens/sContextSource
Llama 3.1 405B969 tokens/sPer user, single queryCerebras press release Jan 2026 [vendor claim]
Llama 3.1 405B2,522 tokens/sUnspecifiedComparison vs NVIDIA DGX B200 (1,038 t/s) [vendor claim]
Llama 4 Maverick2,500 tokens/sPer userCerebras blog 2026 [vendor claim]
Llama4Scout~2,600 tokens/sPer userCommunity tests (Adam Holter, 2026) [independent source, methodology unpublished]

Triangulation: Cerebras does not participate in official MLPerf Inference benchmarks — standardized comparison doesn’t exist. The 969 tokens/s and 2,522 tokens/s figures are Cerebras communications. The only partial triangulation available comes from community tests (Adam Holter, 2026) confirming ~2,600 tokens/s on Llama4Scout via the public Cerebras API — independent source but unpublished methodology.

Architectural limit: the 44 GB of on-chip SRAM isn’t sufficient for very large models (Llama 3.1 405B ≈ 800 GB in FP16). Cerebras uses a disaggregated inference architecture: preprocessing on AWS Trainium for prefill, then the WSE-3 handles decoding. [source: AWS blog March 2026, SiliconAngle March 2026]

2025-2026 trajectory

  • January 2026: OpenAI signs a multi-year partnership for 750 MW of Cerebras capacity, contract estimated at more than $10 billion. [source: SiliconAngle Jan 2026, confirmed Axios Apr 2026 — 3 independent sources]
  • March 2026: AWS deploys the WSE-3 in its datacenters via Amazon Bedrock — first hyperscaler integration. [source: SiliconAngle March 2026]
  • April 2026: Cerebras files its S-1 for a Nasdaq IPO (ticker CBRS), targeting a valuation of at least $35 billion, with 2025 revenues of $510 million. [source: TechCrunch Apr 2026, TheAIInsider Apr 2026]

Groq LPU — the acquisition that changes everything

Deterministic architecture: zero runtime decisions

Groq invented the LPU (Language Processing Unit), a fundamental architectural break from GPUs. Where a GPU makes decisions at every cycle (hardware scheduler managing resource contention, cache misses, branch prediction), the LPU is entirely compiled and deterministic: the Groq compiler (GXCC) pre-computes at compile time the complete execution graph, including inter-chip communications, down to the individual clock cycle. At runtime, the hardware decides nothing.

Vocabulary

  • Jitter: variability in response time. A system with 100 ms average latency and 50 ms jitter will sometimes respond in 50 ms, sometimes in 150 ms. Jitter is critical for interactive applications — users feel latency spikes.
  • Tail latency: latency at the 99th or 99.9th percentile. Often 5-10× the median on GPU, due to resource contention. The LPU eliminates this by design.
  • SRAM: on-chip memory, fast but costly and limited in quantity. See HBM for external GPU memory.

Internal LPU architecture:

  • SRAM only as primary working storage (no HBM or hierarchical cache)
  • SRAM bandwidth: over 80 TB/s vs ~8 TB/s for H100 HBM [source: Groq architecture whitepaper 2024]
  • Specialized execution modules: Matrix (dense tensors), Vector (pointwise arithmetic), Switch (structured data movement)

In plain terms: a GPU improvises at runtime like a chef adapting a recipe based on what’s in the fridge. The LPU executes a score written down to the last note — zero improvisation, zero unexpected delay, but the program must be complete before starting.

Pre-acquisition benchmarks (independent data)

Artificial Analysis (2024) data constitutes the most reliable source available outside MLPerf for Groq:

ModelTokens/sSource
Llama 3 8B877 tokens/sArtificial Analysis (independent, 2024) — triangulated
Llama 3 70B284 tokens/sArtificial Analysis (independent, 2024) — triangulated
Llama 2 70B~300 tokens/sGroq + EE Times 2024 — partially triangulated

These figures compare to an H100 single-GPU achieving approximately 24 tokens/s decoding Llama 3.1 70B (single user) — Groq shows a ~12× ratio on this model. This ratio is consistent with the architectural SRAM vs HBM advantage.

The end of independent Groq: Nvidia acquisition December 2025

This is the major disruption of 2025. Nvidia acquired Groq on December 24, 2025 for $20 billion (structure: non-exclusive license + talent acquisition + IP). Jonathan Ross (founder), Sunny Madra (president), and the majority of engineers joined Nvidia. [source: Tom’s Hardware Jan 2026, IntuitionLabs Jan 2026]

The concrete result: Groq 3 LPX — the LPU integrated into Nvidia’s Vera Rubin platform as an inference co-processor. LPX rack specs:

  • 315 PFLOPS FP8 (256 Groq 3 LPU chips per rack)
  • 128 GB total SRAM
  • 40 PB/s on-chip SRAM bandwidth
  • AFD (Attention-FFN Disaggregation) architecture: Vera Rubin GPU handles prefill and attention; LPU handles decode, FFN, and MoE
  • Nvidia claim: “35× higher inference throughput per megawatt” vs competing solutions [vendor claim, not independently verified]

[source: NVIDIA Technical Blog “Inside Nvidia Groq 3 LPX”, April 2026]

In plain terms: Groq is no longer an independent challenger. The LPU technology is now a building block in Nvidia’s platform — which reinforces Nvidia’s dominance over inference rather than challenging it. Former Groq Cloud customers can now access the technology only through the Nvidia ecosystem.


SambaNova SN40L/SN50 — reconfigurable dataflow for very large models

The RDU: compiling complete graphs, not individual kernels

SambaNova’s approach differs from the two previous ones. The RDU (Reconfigurable Dataflow Unit) doesn’t just accelerate individual operations (matmul, attention) — it compiles complete operation graphs into a continuous dataflow pipeline. Where a GPU runs a sequence of separate kernels (matmul, then softmax, then layernorm), the RDU fuses hundreds of operations into a single call.

SN40L (2023-2025, TSMC 5nm, 2.5D CoWoS chiplet):

  • 1,040 Pattern Compute Units (PCUs): 638 BF16 TFLOPS peak
  • Three-tier memory system:
    • PMU on-chip SRAM: 520 MiB at hundreds of TB/s
    • Co-packaged HBM: 64 GiB at ~2 TB/s
    • DDR DRAM: up to 1.5 TiB at over 200 GB/s (total capacity for very large models)

[source: arxiv 2405.07518 “SambaNova SN40L: Scaling the AI Memory Wall”, SambaNova engineers’ paper presented at Hot Chips 2024]

The 1.5 TiB DDR memory is the distinctive element. It allows multiple very large models to coexist in memory and switch between them 31 times faster than an equivalent GPU system — critical for CoE (Composition of Experts) deployments where several specialized models alternate based on the request.

In plain terms: imagine a library with three zones — the most-used books on your desk (SRAM), classics on nearby shelves (HBM), and archives in an adjacent room (DDR). The SN40L can access the archives without clearing the desk first — a decisive advantage for models with over a billion parameters.

SN40L benchmarks

ModelMetricValueSource
Llama 3.1 405Btokens/s per user114 tokens/sArtificial Analysis (independent) — triangulated
DeepSeek-R1 671Btokens/s per user198 tokens/sTechRadar Jan 2026 (citing SambaNova) — partially triangulated
DeepSeek (aggregate)total throughput>4,500 tokens/sSambaNova blog [vendor claim]

The most striking figure: 198 tokens/s on DeepSeek-R1 671B with only 16 SN40L chips. An equivalent GPU system would need approximately 130 H100s for this 671-billion-parameter model. [source: TechRadar Jan 2026]

SN50 — next generation

Announced in February 2026, with an additional $350 million raised:

  • TSMC 3nm process, dual-chiplet design
  • 3.2 PFLOPS FP8 and 1.6 PFLOPS BF16 per chip
  • 432 MB on-chip SRAM / 64 GB HBM2E / up to 2 TB DDR5
  • Vendor claim: Llama 3.3 70B at 895 tokens/s per user (FP8) vs 184 tokens/s on Nvidia B200 — 4.9× ratio [vendor claim — chip undelivered]
  • Intel partnership for Xeon CPU platform

Expected delivery: second half of 2026 — announced, not delivered at time of writing. SN50 figures remain entirely vendor claims.

[source: BusinessWire Feb 2026, Tom’s Hardware Feb 2026, NextPlatform Feb 2026]


Etched Sohu — the existential bet on transformer

Transformer-only hardwired in silicon

Etched is the most radical bet in this comparison. Where Cerebras uses an entire silicon wafer but remains a relatively general processor, Sohu is a pure transformer-only ASIC: multi-head attention blocks (QKV projections, softmax, output projection), feed-forward networks (GELU/SiLU), and layer normalization are hardwired in silicon. Sohu cannot be reprogrammed for a different model type.

Theoretical advantage: maximum transformer compute density, no surface wasted on unused operations. Announced specifications (TSMC 4nm):

  • 144 GB HBM3E
  • Claim: 500,000 tokens/s on Llama 70B for an 8-Sohu server [vendor claim]
  • “Replaces 160 H100s” according to Etched communications [vendor claim]

[source: The Register June 2024, Tom’s Hardware June 2024]

In plain terms: a Sohu is like a professional espresso machine that makes exactly one type of coffee perfectly — but cannot make tea, heat food, or do anything else. If espresso remains the standard, it’s a genius machine. If habits change, it’s useless hardware.

The structural disadvantage: the existential bet

The transformer architecture has dominated since 2017, but that’s a convention, not a physical law. If the architecture evolves significantly (post-transformer, State Space Models, Mamba-style, hybrid), Sohu becomes useless. This is an existential bet on the permanence of transformers.

Funding and actual status (April 2026)

EventDateAmountSource
Series AJune 2024$120MThe Register + Crunchbase — triangulated
Additional fundingJanuary 2026$500MData Center Dynamics Jan 2026
Total raised>$620MValuation: $5 billion

Delivery status: zero customer deliveries as of April 2026. A Manifold Markets prediction (“Will Sohu ship within a year of announcement?”) resolved NO in mid-2025. As of March 2026, no external customer had received a Sohu. Etched mentions “tens of millions of dollars in reservations” from unnamed customers.

Zero independent benchmarks: all cited performance figures (500,000 tokens/s, “20× vs H100”) are theoretical projections from internal simulations. No third party has tested Sohu under real conditions. [source: Tom’s Hardware “Sohu AI chip claimed to run models 20x faster”, Manifold Markets prediction, Data Center Dynamics Jan 2026]

In plain terms: with $620 million raised and no delivery or third-party benchmark, Etched is the hardest case to evaluate. The architecture could be brilliant or catastrophic — there’s no data yet to know. Treat all announced performance as [unverified — no third-party data].


Tenstorrent Blackhole — the open-source approach

Multi-market strategy and open source

Tenstorrent is the most atypical player in this comparison. Its strategy, led by Jim Keller (former architect at Intel, AMD, Apple, Tesla), targets multiple markets simultaneously: datacenter inference, edge and embedded inference, and RISC-V IP licensing to other manufacturers. The software stack (tt-metal) is fully open-source — a rarity in the specialized ASIC ecosystem.

Tensix architecture: RISC-V + matmul

The core of Tenstorrent is the Tensix core. Each Tensix integrates a 5-core RISC-V CPU (for orchestration) plus vector units and a matrix engine. On Blackhole (2024, TSMC 6nm):

  • 120 Tensix cores per chip (note: reduced from 140 by firmware update, Tom’s Hardware Nov 2025)
  • 32 GB GDDR6 per chip (not HBM — an economic choice)
  • 210 MB SRAM distributed on-chip
  • 774 TFLOPS FP8 (Blackhole p150) [source: Tenstorrent.com/hardware/blackhole]

The previous generation, Wormhole n300 (still available for purchase), integrates 128 Tensix cores, 192 MB SRAM, and 24 GB GDDR6.

[source: AnandTech Wormhole launch 2024, Tenstorrent.com/hardware/blackhole]

Real benchmarks: 50% of theoretical peak

The Register published in November 2025 an independent test of the QuietBox — a Tenstorrent Blackhole workstation sold for approximately $12,000:

  • Real performance: ~50% of theoretical peak on Llama 3.1 8B
  • 76 of the 120 Tensix cores (before correction) unused during LLM inference — kernels optimized for Wormhole have not yet been ported to Blackhole
  • Single P150 performance ≈ Nvidia DGX Spark on tested models, despite a theoretical 2-3× advantage for Blackhole
  • The Register verdict: “the lack of optimized kernels for LLM inference is an inexcusable shortcoming”

[source: The Register “Blackhole QuietBox review” Nov 2025 — independent source]

In plain terms: Tenstorrent delivers hardware, but the software to exploit that hardware isn’t yet at the right level. It’s like having a Formula 1 engine with bicycle tires — the power is there, but it can’t express itself. Important nuance: tt-metal is actively developed on GitHub, and November 2025 figures are not the final word.

Funding and trajectory

  • $693 million raised (round led by AFW Partners and Samsung Securities, with LG, Fidelity, and Bezos Expeditions) — $2.6 billion valuation
  • $150 million in signed contracts at the time of the raise
  • November 2025: discussions for an additional $800 million round at $3.2 billion valuation (Fidelity lead)
  • Roadmap: bi-annual hardware cycle (Blackhole 2024 → next generation planned ~2026)

[triangulated: Yahoo Finance, Tom’s Hardware, TechSpot, TheStack — 4 independent sources on the $700M round]


What the market did with these bets — September 2026 review

A year on, the question is no longer “do these architectures work?” but “who is still independent?”. The answer is instructive.

PlayerWhat moved in 2026
GroqAbsorbed. The Nvidia acquisition, $20 billion, closes in December 2025.
CerebrasWent public in May 2026 — $5.6 billion raised, the year’s largest IPO, initial market capitalisation near $40 billion. Three months earlier, a $1 billion Series H at a $23 billion valuation, closed four days after signing a $10 billion multi-year compute agreement with OpenAI.
SambaNovaA $1 billion first close at roughly $11 billion post-money valuation in July 2026. The SN50 is due to ship to customers by the end of the year.
EtchedMore than $1 billion in signed customer contracts after demonstrating working silicon — the racks are still under validation, the terms remain private.
TenstorrentAn order of 96 Galaxy units, that is 3,072 Blackhole chips, reported by Jim Keller.

These figures come from the trade press and company statements; none is audited. [UNVERIFIED]

In short: the non-GPU bet did not fail, it sorted itself out. The most technically accomplished — Groq and its cycle-exact determinism — was bought by the company it was meant to compete with. The two that survived independent did so by leaning on a single massive customer (OpenAI for Cerebras) or by raising without interruption. The only player in the table selling a product a reader can buy and put on a desk is still Tenstorrent.


What independent benchmarks actually show

There is a structural limitation in any comparison of these players: none of them regularly participate in MLPerf Inference benchmarks (the industry’s official benchmarks, arbitrated by an independent third party). Standardized comparison doesn’t exist. Available data comes from three sources of unequal quality:

Most reliable data — Artificial Analysis (independent cloud benchmark service):

Chip / ModelTokens/s per userStatus
Groq LPU / Llama 3 8B877Triangulated — independent source
Groq LPU / Llama 3 70B284Triangulated — independent source
SambaNova SN40L / Llama 3.1 405B114Triangulated — independent source
H100 GPU single / Llama 3.1 70B~24Established reference

Partially triangulated data (1-2 independent sources, partial methodology):

  • SambaNova 198 tokens/s on DeepSeek-R1 671B with 16 SN40L chips (TechRadar Jan 2026)
  • Cerebras ~2,600 tokens/s on Llama4Scout (Adam Holter, community, 2026)

Vendor claims only (not verified by third parties):

  • Cerebras 2,522 tokens/s Llama 3.1 405B vs DGX B200 (1,038 t/s)
  • Etched Sohu 500,000 tokens/s Llama 70B (8-chip server)
  • SambaNova SN50 895 tokens/s per user Llama 3.3 70B
  • Nvidia Groq 3 LPX “35× inference throughput/MW”

In plain terms: without MLPerf, the only rigorous comparative baseline available in April 2026 is Artificial Analysis for Groq and SambaNova. All other figures — including the most spectacular ones — should be treated with caution.


Decision matrix — which technology for which use case

ContextActor to considerWhy
Ultra-low conversational latency, models >100B parametersCerebras WSE-3On-chip SRAM eliminates memory bottleneck; OpenAI + AWS partnerships validate production use
Deterministic latency, zero jitter, real-time applicationsGroq (via Nvidia)LPU eliminates runtime decisions by design; independent data confirms 877 t/s on 8B
Very large MoE/multilingual models (>500B parameters), rapid model switchingSambaNova SN40L1.5 TiB DDR memory + 31× faster switching; Artificial Analysis data on 405B triangulated
Hardware developers, edge enterprise, open sourceTenstorrent BlackholeOpen-source tt-metal stack, ~$12K workstations, RISC-V licensing; but only 50% of LLM peak as of Nov 2025
Bet on pure transformer-only inference at very high density (H2 2026+)Etched SohuZero terrain data available; $620M raised without delivery; consider only if transformers remain dominant paradigm at 3-5 year horizon

Decision heuristics

Never compare tokens/s without specifying context. “1,000 tokens/s” aggregated over 50 simultaneous requests corresponds to 20 tokens/s per user — mediocre conversational speed. Always ask: aggregate or per-user tokens/s? Which model? Which batch size? Which quantization?

Vendor claims must be triangulated. None of these five players participate in MLPerf. The only rigorously comparable data available in 2026 are those from Artificial Analysis (Groq, SambaNova) and community tests for Cerebras. Everything else is unverified marketing.

The ecosystem matters as much as the chip. Groq, the most credible challenger on the latency metric, was absorbed by Nvidia. SambaNova and Cerebras are integrating into AWS. Tenstorrent bets on open source. The chip alone isn’t enough.

Distinguish delivered vs. announced. In April 2026: WSE-3 delivered and in production at AWS + OpenAI. SN40L delivered. Groq LPU delivered (but now Nvidia). Tenstorrent Blackhole delivered (partial software maturity). Sohu: zero customer deliveries.


What to take away in 2026

The AI inference ASIC startup market is going through a forced clarification in 2026. Three signals:

First signal — validation vs. promises. The gap between players validated in production (Cerebras with OpenAI + AWS, SambaNova with DeepSeek + enterprise customers) and unvalidated players (Etched, still Tenstorrent on pure inference) has widened. An ASIC remains a promise as long as no external customer has tested it in real production.

Second signal — consolidation has begun. Nvidia’s $20 billion acquisition of Groq is not a detail — it’s a signal that LPU technology was valid enough for Nvidia to want to integrate it into its Vera Rubin platform. And it’s simultaneously proof that surviving independently against Nvidia is difficult even with the industry’s best latency metrics.

Third signal — the question that matters in 2026. Blackwell GPUs (H200, B200, B300) are improving inference efficiency much faster than expected. The real question is no longer “who makes the fastest token?” but: “Will GPU inference cost deflation happen fast enough for hyperscalers to remain customers of these startups, or will they all internalize as AWS (Trainium), Google (TPU), and Meta (MTIA) have already done?”

The five architectures described here provide different answers — but the market’s response will be read in 2026-2027 revenues, not in synthetic benchmarks.