Skip to content

Infrastructure

The hardware, cloud and energy powering LLMs

16 articles · 4 sous-catégories
Subcategory
Content type
comparatif GPUs & chips

AI Inference ASIC Alternatives 2026 — Five Startups Rethinking the Chip for LLMs

Cerebras, Groq (absorbed by Nvidia), SambaNova, Etched, Tenstorrent: radically different architectures, contrasting benchmarks, and one shared lesson — the chip alone isn't enough, it's the ecosystem that wins.

cerebrasgroqsambanovaetchedtenstorrentasicinferencelpuwafer-scalerdublackholewse-3sn40lsohuhardwarellm-inferencelatencytokens-per-second
concept Datacenters & energy

Liquid cooling — how AI datacenters dissipate 120 kW per rack

Water carries 3,500 times more energy than air per unit volume. When a GPU rack exceeds 40 kW, the laws of physics force a shift to liquid. A technical guide to DLC, immersion and rear-door techniques, PUE/WUE metrics, and real operational trade-offs.

coolingdlcimmersionpuewuenvl72ashraeb200rear-doorrdhxcold-platecdutwo-phasesingle-phaseblackwell
concept Memory & storage

CXL and memory pooling — sharing memory across servers

CXL (Compute Express Link) allows multiple servers to share a common memory pool over PCIe — a technology distinct from GPU HBM and compute networking, that changes the equation for KV cache in long-context LLM inference.

cxlmemory-poolingpciesapphire-rapidsgenoakv-cachememory-expandermemory-disaggregationdramhbmllm-inferencelong-contexttiered-memory
analyse Economics

AI Economics — Why Inference Now Costs More Than Training

From GPT-3 to o3, the real economics of generative AI have flipped: training costs tens to hundreds of millions of dollars, but inference — billions of cumulative requests — devours the budgets. A data-driven analysis of the CAPEX → OPEX shift, negative margins, and the pricing war.

inferencetrainingcostscapexopexreasoning-modelsapimarginsopenaianthropicdeepseekgpupricing
concept Memory & storage

CPO and optical interconnects — when light replaces copper in AI datacenters

AI clusters are hitting a physical wall: copper can no longer carry the bandwidth required between GPUs. Co-Packaged Optics (CPO) integrates photonic components directly into chip packages to cut interconnect power consumption by more than ten times.

cposilicon-photonicsayar-labslightmattercelestial-aibroadcomquantum-xtsmc-coupeinterconnectsdatacenteropticalpj-bit
analyse GPUs & chips

Nvidia Roadmap 2025-2028 — From Blackwell Ultra to Feynman, Four Generations at Annual Cadence

Nvidia adopted an annual release cadence for data center GPUs in 2024. This article details the four upcoming generations — Blackwell Ultra (B300), Vera Rubin (R100), Rubin Ultra (R300), and Feynman — with technical specs, HBM4 supply constraints, and datacenter implications.

nvidiarubinvera-rubinfeynmanblackwell-ultrab300r100r300hbm4nvlinktsmcroadmapdatacentergpusk-hynixsamsung
analyse Cloud & sovereignty

European Sovereign Cloud — when hosting becomes a matter of law

In Europe, hosting data for AI is no longer enough: you must also choose an operator not subject to US law. An overview of the players, the rules, and the real trade-offs of sovereign cloud in 2026.

sovereign-cloudovhcloudscalewayoutscalegaia-xai-actcloud-actsecnumcloudeucst-systemsorange-businessbleus3nsmistral-airgpd
comparatif GPUs & chips

TPU vs GPU in 2026 — the real Ironwood versus Blackwell face-off

For the first time since the H100, Google and Nvidia are launching their flagship chips in the same time window. An in-depth technical comparison: architectures, MLPerf v4.1/v5.0 benchmarks, software ecosystems, use cases and 2026 costs.

tpugpuironwoodblackwellgb200nvl72mlperfjaxcudajax-xlabenchmarktpu-v7b200inferencetraining
concept Datacenters & energy

Datacenters and energy — what AI actually consumes

Generative AI has multiplied compute density in datacenters tenfold. Energy, water, carbon: real orders of magnitude, open controversies, and hyperscaler strategies.

datacenterenergygpucoolingcarbon-footprintwatergreen-ai
concept GPUs & chips

Your phone has an AI chip — here's what it can do

Running an LLM on a smartphone is possible. But the hardware, frameworks, and compact models impose trade-offs that fundamentally change what a model can do.

edge-aiinferencemobileon-devicequantizationlatency
concept GPUs & chips

Nvidia GPUs and CUDA — why they dominate AI compute

Nvidia GPUs have become the base infrastructure of large language models. Understanding why requires explaining CUDA as much as the hardware itself.

gpunvidiacudah100blackwellinferencedatacenter
concept Memory & storage

High Bandwidth Memory — When Memory Becomes the GPU Bottleneck

The most powerful GPUs are often waiting, not computing. HBM, ultra-fast memory bonded directly to accelerators, has become the true limiting factor in AI.

hbmgpumemoryinferencememory-wallbandwidth
concept Datacenters & energy

High-Performance Networking — The Invisible Backbone of AI

GPUs make the headlines, but the network between them determines whether a 10,000-chip cluster runs at 50% or 5% of its capacity.

networkinginfinibandrdmadatacenternvidiahigh-performance
concept GPUs & chips

Alternatives to Nvidia — Who Is Really Challenging the AI Chip Market?

AMD, AWS, Cerebras, Tenstorrent: a survey of challengers seeking to reduce dependence on Nvidia in large model training and inference infrastructure.

chipsamdintelgroqcerebrasalternativeshardware
concept Supply chain

Semiconductor supply chain — why your GPU comes from everywhere but where you live

Every AI chip crosses multiple countries and technical monopolies before landing in a datacenter. A look at the geographic and technological dependencies that govern GPU access.

supply-chainsemiconductorstsmcasmlgeopoliticsexport-controls
concept Datacenters & energy

AI and the environment — energy, water, carbon

Training a large language model can emit as much CO₂ as several cars over their lifetime. But the real debate is about inference, transparency, and the rebound effect.

energycarbonwaterdata centersfootprintgreen AIinferencetraining