In brief

In 2024, a rack housing 72 B200 GPUs (Nvidia’s NVL72 configuration) dissipates 120 kilowatts continuously — the equivalent of 60 electric stoves running at full power inside a 60 cm × 120 cm footprint. Air simply cannot remove that heat. Water can: its thermal capacity is 3,500 times greater than air’s per unit volume. This physical fact explains why hyperscalers (Microsoft, Meta, Google, CoreWeave) are switching massively to liquid cooling — and why this shift is irreversible for high-density AI workloads.

This article covers the concrete thermal mechanics: why air fails beyond a threshold, how the four major cooling families work (air, rear-door, direct-to-chip, single- and two-phase immersion), how efficiency is measured with PUE and WUE metrics, and what the real operational risks are. It does not cover global datacenter power consumption (→ Datacenters and energy) or the lifetime carbon and water footprint of AI systems (→ AI and the environment).


Understanding the problem: when physics sets the rules

The power density trajectory

For fifteen years, datacenter teams optimized the same system: blow cold air under raised floors, let heat rise, capture it in a hot aisle. In 2010, a standard rack dissipated 4 to 10 kW. By 2020, H100 DGX servers pushed into the 20-40 kW range. In 2024, the GB200 NVL72 rack reaches 120-140 kW.

PeriodTypical power per rackThermal regime
2010-20154-10 kWAir cooling sufficient
2015-202010-20 kWOptimized air (hot/cold aisles)
2020-202320-40 kW (H100 DGX)Tipping point toward liquid
2024-2026120-140 kW (GB200 NVL72, Blackwell Ultra)Liquid mandatory

Source: Tom’s Hardware, “The data center cooling state of play (2025)”.

This density progression is not a choice — it is driven by Moore’s Law applied to AI accelerators: each new GPU generation packs more transistors into the same package and dissipates more watts. The NVL72 rack concentrates 72 B200 GPUs and 36 Grace CPUs into approximately 42 rack units.

The physical limits of air

Imagine trying to cool a forge with a desk fan. Air can carry some heat, but not fast enough for such an intense source. That is precisely the situation facing high-density AI racks against forced convection cooling.

The cause is measurable: the volumetric heat capacity of air at 20°C is approximately 1,200 joules per liter per kelvin. Water’s is 4,180,000 joules under the same conditions — a ratio of 3,483.

In plain terms: to carry the same amount of heat, air must flow 3,500 times faster than water, or in 3,500 times the volume. At 120 kW per rack, the required air velocities exceed 15 m/s — incompatible with any enclosed building infrastructure, and destructive to the hardware itself.

NVL72 concrete case: NVIDIA mandates a liquid cooling architecture with an external Coolant Distribution Unit (CDU) sized at 150-200 kW, a minimum flow rate of 2-3 liters per minute per module (approximately 80 liters per minute for the full rack), and an inlet liquid temperature between 20 and 45°C. (Source: ToneCooling, NVIDIA GB200 NVL72 Cooling Requirements, 2024.)

The market shift

This physical determinism translates into market figures: liquid-based systems represent 46% of the global datacenter cooling market in 2024, up from approximately 10% in 2020. The datacenter liquid cooling market grew from $1.5 billion in 2024 toward a projected $7 billion by 2029. (Sources: Tom’s Hardware 2025; Dell’Oro Group, 2024.)


The four cooling families

There is a continuum of solutions, from simplest to most sophisticated. Each family targets a density range and comes with its own infrastructure constraints.

Core vocabulary

  • Forced convection: moving a fluid (air or liquid) to transfer heat from one point to another.
  • Volumetric heat capacity: the amount of energy a fluid can absorb per unit of volume and temperature difference. Higher means more effective as a heat carrier.
  • CDU (Coolant Distribution Unit): the interface between the primary cooling circuit (server side) and the secondary circuit (building/chiller side).
  • PUE (Power Usage Effectiveness): total datacenter energy / IT equipment energy. Perfect theoretical value = 1.0.
  • WUE (Water Usage Effectiveness): liters of water consumed per kilowatthour of IT energy.

Family 1 — Air cooling (< 40 kW per rack)

Air cooling is the historical baseline. The standard infrastructure: CRAC (Computer Room Air Conditioner) or CRAH (Computer Room Air Handler) units blow cold air under a raised floor, servers heat it, hot air rises toward the hot aisle. Free cooling — air-to-air heat exchangers using cold outdoor air — significantly improves efficiency in temperate regions.

Performance: achievable PUE between 1.2 and 1.4 in favorable climates, 1.5 to 1.8 in hot regions without free cooling.

Limit: physically inadequate beyond 40-50 kW per rack in dense clusters. Still appropriate for standard servers (below 20 kW per rack) or in mixed configurations within the same facility.

In plain terms: air cooling remains valid for office workloads and standard application servers. The moment you install a current-generation GPU rack, you need something else.

Family 2 — Rear Door Heat Exchanger (30-60 kW per rack)

Picture a central heating radiator mounted at the back of a network cabinet: that is exactly the principle of the Rear Door Heat Exchanger (RDHx). Hot air exiting the servers passes through a water-to-air heat exchanger before being discharged into the room. Two variants exist:

  • Passive RDHx: no dedicated fans — uses only the servers’ internal fans. For moderate loads (up to ~30 kW per rack depending on server airflow).
  • Active RDHx: dedicated fans integrated into the door. Improves thermal throughput, suited for higher loads (up to 50-60 kW per rack).

Performance: 30 to 60% more efficient than air alone, according to Silverback Data Center Solutions. Key advantage: installs in an existing datacenter without running pipes to every rack location.

Adoption: 29% of datacenter operators already use RDHx; 83% of new projects plan deployment within the year. (Source: InsightAce Analytic, Data Center RDHx Cooling Market 2026-2035 [commercial market estimate].)

Limit: a transition solution, not an endpoint. Passive RDHx hits its limits around 30 kW per rack; active around 60 kW. Beyond that, only Direct-to-Chip or immersion cooling suffices.

In plain terms: RDHx is the path of least resistance for operators who need to deploy moderately dense GPUs into existing infrastructure. It is not the final answer for NVL72-class racks.

Family 3 — Direct-to-Chip Liquid Cooling / DLC (50-100+ kW per rack)

Rather than cooling the room air, DLC brings the liquid directly to the heat source — the chip itself. Cold plates made of copper or aluminum are mounted directly on CPUs and GPUs. A coolant (deionized water, water-glycol mix, or propylene glycol) flows through microchannels machined into the plate, absorbs heat within a few millimeters of the silicon, then returns hot to the CDU.

NVL72 reference architecture:

  • 36 cold plates mounted in parallel on a common manifold (supply/return circuit)
  • Total flow rate: approximately 80 liters per minute for the complete rack
  • Tolerated inlet temperature: up to 45°C (B200 GPUs accept warm inlet fluid, reducing the need for energy-intensive mechanical chillers)
  • External CDU sized at 150-200 kW

Key advantage: DLC does not require server redesign. Cold plates adapt to existing hardware architectures. Residual air remains for peripheral components — storage, network interfaces, power supplies — that lack cold plates.

Adoption: 47% share within the AI liquid cooling segment in 2024 (DLC is the dominant segment). 22% of datacenter operators already use DLC; 61% are considering it. (Source: Uptime Institute, Cooling Systems Survey 2024.)

Limit: requires complete hydraulic infrastructure (piping, fittings, CDUs). Leakage risk — water or glycol near energized equipment — is the primary operational concern. ASHRAE TC 9.9’s September 2024 technical bulletin explicitly states that “loss of cooling can be catastrophic when supporting extreme chip powers” and recommends explicit resilience protocols.

In plain terms: DLC is the de facto technology for high-density GPU deployments in 2025-2026. It strikes the best balance between high thermal efficiency and infrastructure compatibility with existing datacenter practices.

Family 4 — Immersion cooling (100+ kW per rack)

In immersion cooling, there are no cold plates or local pipes: servers are fully submerged in a bath of dielectric fluid. Two sub-families exist with very different properties.

4a — Single-phase immersion

The fluid (mineral oil or synthetic fluid — e.g., Shell Immersion Cooling Fluid) remains liquid, absorbs heat by flowing around components, and is pumped to an external heat exchanger.

Advantages: complete elimination of server fans (measured savings of ~80% in fan power consumption), thermal uniformity across the entire server, densities up to 200+ kW per rack theoretically supported.

Real deployments: Meta showcased 140 kW liquid-cooled racks for LLaMA workloads at the OCP Global Summit in October 2024. The global immersion cooling market is estimated at $5.72 billion for 2026. (Source: Mordor Intelligence, 2025.)

Limits: servers must be extracted from standard racks and placed in sealed pods — maintenance is more demanding. Lock-in on dielectric fluids (often proprietary) and higher cost than DLC slow adoption.

4b — Two-phase immersion

More sophisticated variant: the dielectric fluid (historically fluorocarbon-based, such as 3M Novec) changes state from liquid to vapor on contact with hot chips. The latent heat of vaporization absorbs thermal energy far more efficiently than a stable liquid. Vapor rises, condenses on a cooling coil at the top of the tank, and falls back as liquid by gravity — an autonomous cycle requiring no pump for the gas phase.

Performance: best heat transfer efficiency of all techniques. Densities above 200 kW per rack theoretically supported.

Deployments: Microsoft tested two-phase immersion on AI training clusters with a reported 30% energy efficiency gain [unverified — secondary sources only, original Microsoft publication not located]. Meta is deploying two-phase immersion in 2024-2025.

Critical limits: fluorocarbon fluids face strong regulatory pressure (PFAS, high global warming potential). 3M announced a phased discontinuation of Novec products by end 2025, creating supply tension for two-phase fluids. Complete rack redesign is mandatory — no transition from air or DLC without breaking server mechanics.

In plain terms: two-phase immersion is the most thermally efficient technology, but also the most complex to deploy and the most exposed to regulatory risks on fluids. It remains a high-growth niche, not yet the hyperscale standard.


Efficiency metrics: PUE and WUE

Comparing cooling installations without shared metrics is like comparing cars without fuel consumption figures. The industry has developed two standardized indicators.

PUE — Power Usage Effectiveness

PUE measures the overall power efficiency of a datacenter. A PUE of 1.0 means all consumed energy goes to computation; none is wasted on cooling or infrastructure. In practice, cooling represents the main gap between 1.0 and the actual PUE.

Installation typeTypical PUE
Classic datacenter (global average 2024)1.56
Optimized air-cooled datacenter1.2-1.4
Liquid-cooled datacenter (DLC/immersion)< 1.2
Google — fleet worldwide (2024)1.09
New AI-first hyperscale facilities1.1-1.15

Sources: Google PUE — Google Data Centers Efficiency page, 2024. Global average — Lawrence Berkeley National Laboratory, 2024.

Measured impact of switching to liquid (Vertiv study on a Tier 2 datacenter, ~1.5 MW, Baltimore region): moving from 100% air to 75% liquid cooling reduces installation power by 18.1% and total datacenter power by 10.2%. PUE drops from 1.38 to 1.34 — a modest ratio gain, but significant absolute savings. (Source: Vertiv, Quantifying the impact on PUE when introducing liquid cooling, EMEA.)

Geographic effect: PUE depends strongly on climate. A datacenter in Montreal with free cooling available seven months per year can reach PUE 1.10-1.15. A datacenter in Phoenix without liquid cooling struggles below PUE 1.35. With DLC at high inlet temperature (up to 45°C for B200 GPUs), heat rejection can be achieved with less energy-intensive cooling towers, even in hot climates.

In plain terms: PUE drops toward 1.0 as liquid cooling replaces air. Every PUE point gained directly represents less electricity wasted on thermal infrastructure. For a 100 MW datacenter, moving from 1.4 to 1.2 saves 20 MW of permanent power draw.

WUE — Water Usage Effectiveness

WUE measures datacenter water consumption in liters per kilowatthour of IT energy. Developed by The Green Grid, it specifically targets evaporative water used in cooling towers — a tension point between liquid cooling and environmental impact.

ScenarioWUE
Industry average1.9 L/kWh
Microsoft FY20240.30 L/kWh
Microsoft new design target (zero-evaporation)~0 (building services residual only)
Berkeley Lab 2028 projection (liquid adoption)0.45-0.48 L/kWh

Sources: Microsoft — Sustainable by design blog, December 2024. Berkeley Lab — 2024 US Data Center Energy Usage Report.

Important nuance: liquid cooling does not mechanically reduce water consumption. A closed-loop DLC circuit, where the CDU rejects heat via an air-cooled mechanical chiller, consumes near-zero water (WUE ≈ 0). But if the CDU rejects heat via an evaporative cooling tower, WUE remains high — simply relocated from the server floor to the tower. The actual reduction depends on the complete secondary circuit architecture.

Microsoft zero-water design (August 2024): since August 2024, all new Microsoft datacenter projects integrate a fully closed loop without evaporation. Estimated savings: more than 125 million liters of water per year per datacenter. Microsoft’s WUE reached 0.30 L/kWh in FY2024, a 39% improvement versus 2021. Pilot deployments are underway in Arizona and Wisconsin, with full deployment projected by end 2027. (Source: Microsoft Cloud Blog, December 2024.)

In plain terms: WUE measures where the water goes — a closed-loop DLC system without an evaporative tower can approach WUE = 0. But the same infrastructure connected to a wet cooling tower doesn’t save water; it just moves the consumption. The lifecycle environmental angle is covered in AI and the environment.


Deployment status 2024-2026

Hyperscalers — the transition is underway

Microsoft: since August 2024, every new datacenter project integrates zero-water evaporation cooling in a closed loop. Target density: 140 kW per rack. AI supercomputers deployed in 2025 are 100% liquid-cooled. (Source: Microsoft Cloud Blog, December 2024.)

Meta: 140 kW racks presented at OCP Global Summit in October 2024. Declared goal: 100% of Meta’s datacenters on liquid cooling by 2030. The $65 billion AI infrastructure investment announced for 2025 includes large-scale liquid cooling deployment. (Sources: Meta OCP announcements; EnkiAI analysis, 2024.)

Google: TPU racks are on custom liquid cooling. Fleet-wide PUE of 1.09 in 2024 — best in the sector. (Source: Google Data Centers Efficiency, 2024.)

CoreWeave: all new datacenters on liquid cooling since 2025, for GB200 NVL72 racks at 130 kW. [estimate based on public CoreWeave communications 2024-2025]

Market data (with caveats)

IndicatorValueSourceReliability
Liquid share of datacenter cooling market46% in 2024Tom’s Hardware / Global Market InsightsCross-referenced 2 sources
Operators already using DLC22% in 2024Uptime Institute Survey 2024Self-reported, N unspecified
Operators considering DLC61%Uptime Institute Survey 2024Self-reported
Operators using RDHx29%Uptime Institute Survey 2024Self-reported
Hyperscale AI in liquid cooling 202555%Future Market Insights 2025Commercial market estimate
Liquid cooling market 2024-2029$1.5B → $7BDell’Oro Group 2024Market projection

Methodological note: adoption percentages vary widely depending on whether DLC-only, DLC + immersion, or DLC + immersion + RDHx are counted. Market projections (Dell’Oro, Future Market Insights) are commercial estimates, not independent measurements.

Retrofit versus new build

New build: de facto standard for AI hyperscale. Hydraulic infrastructure (piping, CDUs, manifolds) is integrated at design stage, at marginal additional cost versus pure air infrastructure.

Retrofit: the primary barrier. Adapting an existing air-cooled infrastructure to host GB200 racks is estimated at $5 to $10 million per megawatt [single source — Introl, 2025; not confirmed by independent source]. Associated electrical requirements — 480V three-phase circuits, 300A cables — are in industrial-grade territory.

Barriers cited (Uptime Institute Survey 2025, n = 1,033 operators):

  • Lack of CDU connector standardization: 39%
  • High cost: 38%
  • Limited vendor choice: 26%
  • More complex maintenance: 26%
  • Staff training: 16%

In plain terms: for a new facility planned for high-density AI, liquid cooling is cheaper long-term. For existing infrastructure, retrofit is expensive and requires significant reorganization.


Decision matrix — which technique for which context

Context (rack density)Recommended techniqueWhy
< 30 kW — standard application serversOptimized air cooling (CRAC/CRAH + aisles)Thermally sufficient, mature infrastructure, low operating cost.
30-50 kW — recent GPUs without dense clustersPassive or active RDHxFast retrofit in existing infrastructure, no underfloor piping required.
50-100 kW — modern GPU clusters (H100, A100)DLC mandatoryOnly viable solution at this level without massively over-provisioned air.
> 100 kW — NVL72, Blackwell Ultra, next-gen AI racksDLC or immersion, residual air for peripheralsNVIDIA mandates liquid from the NVL72 spec. Immersion for fan elimination and further density gains.
Greenfield hyperscale projectImmersion or integrated DLC from design stageLower marginal cost than retrofit; better PUE; future density flexibility.

Heuristics:

  1. Start by auditing current and future density: a 20 kW rack today may reach 80 kW at the next GPU generation. Design for the next density, not the current one.
  2. Do not confuse PUE ratio with absolute savings: a PUE of 1.1 in a 10 MW datacenter saves more energy than a PUE of 1.05 in a 1 MW facility.
  3. RDHx is not an endpoint: it is a useful bridge, not a target architecture for 2026+ AI workloads.
  4. Before choosing two-phase immersion, verify the availability of 3M Novec replacement fluids — PFAS regulation is a real supply risk over 2-5 years.
  5. Measure WUE, not just PUE: a DLC system connected to an evaporative cooling tower improves PUE but does not reduce water consumption.

Operational risks

Switching to liquid cooling is not without risk. Here are the main ones, with documented mitigations.

Hydraulic risks

Leaks: water or glycol near energized electronic equipment. Risk of short-circuit or accelerated corrosion. Dielectric fluids (immersion) do not conduct electricity — the electrical risk is absent, but degraded fluid can contaminate components. Standard mitigation: certified piping, quick-disconnect fittings with auto-shutoff, localized leak sensors under raised floors and near CDUs.

Corrosion: deionized water can be aggressive toward certain metal alloys. A water quality management program (conductivity, pH, corrosion inhibitors, biocides) is required — a new discipline for teams trained primarily in airside mechanics.

Manifold pressure: NVL72 manifolds operate typically up to 6 bar, with minimum burst specification at 12 bar. Incidents are rare, but a manifold failure on an active rack can be catastrophic.

Maintenance complexity

Extraction during immersion: removing a server from an oil or fluorocarbon bath requires a heavy procedure — personnel protection, partial drainage, component cleaning. Maintenance cadence is more constrained than in an air rack where replacing a drive takes two minutes.

Connector interoperability: there is no universal standard for CDU ↔ rack connectors. The Open Compute Project (OCP) is working on specifications, but adoption remains fragmented. 39% of operators cite this lack of standardization as their primary barrier to adoption. (Source: Uptime Institute, 2025.)

Training: classic datacenter operations teams lack hydraulic engineering background. Training in liquid cooling — water quality management, leak detection, lock-out/tag-out (LOTO) procedures on pressurized circuits — represents a non-trivial investment.

Resilience: ASHRAE TC 9.9’s September 2024 technical bulletin on high-density liquid cooling resilience recommends explicit continuity protocols in case of coolant loss — something air architectures handle naturally through the permanent presence of ambient air.

In plain terms: liquid cooling is not plug-and-play. It introduces mechanical systems (pumps, piping, heat exchangers) with their own failure modes, and requires operational skills that most datacenter teams will need to develop from scratch.