In short
In 2024, datacenters consume 415 TWh of electricity worldwide — approximately 1.5% of total global consumption. The International Energy Agency projects this figure will reach 945 TWh by 2030, a doubling in six years driven largely by artificial intelligence workloads. To meet this demand, major cloud operators are signing nuclear contracts, restarting power plants, and switching their infrastructure to liquid cooling. But several fundamental controversies remain open: actual efficiency, carbon accounting, water consumption.
In short: an AI datacenter is a city-factory. Its electricity consumption is measured in gigawatts (the equivalent of a nuclear power plant), it pumps water from an entire river to cool itself, and it eats up grid capacity in the regional electrical system. At the global scale, these city-factories stack up the equivalent of a new Norway every six years.
Orders of magnitude
The United States concentrates a massive share of this compute: 176 to 183 TWh consumed by American datacenters in 2023–2024, representing 4.4% of national electricity (Lawrence Berkeley National Laboratory, December 2024). By way of comparison, this exceeds the annual residential consumption of the state of Texas.
What has changed is the nature of the workloads. AI model inference and training servers are built around graphics processors (GPUs) whose power consumption bears no comparison with a conventional web server. AI workloads are growing at 30% per year, versus 9% for conventional server workloads (IEA, Energy and AI, 2025). Capital expenditure by the four main hyperscalers — Amazon, Microsoft, Google, Meta — exceeded 200 billion dollars in 2024, with 2025 forecasts between 315 and 400 billion.
The LBNL 2024 report proposes an estimation range for the United States in 2028: between 325 and 580 TWh, or 6.7 to 12% of national electricity. The gap between these two bounds reveals the structural uncertainty weighing on all projections.
Training and inference: two types of cost
Understanding AI’s footprint begins by distinguishing two radically different phases.
Training is the one-time but massive cost of training a model. It mobilizes thousands of GPUs for weeks or months. This is the spectacular figure cited in press articles.
Inference is the cost of each user request. Individually small, but cumulated across billions of daily calls, it represents a growing — sometimes comparable — share of a model’s total footprint over its lifetime.
An article by Luccioni et al. on BLOOM, a 176-billion-parameter model trained on the Jean Zay supercomputer in France, shows that eighteen months of inference produced a carbon footprint comparable to that of training itself (24.7 tCO₂e for training). This is the first academic study to document this balance on a model of this size.
For more recent models, official data is scarce. For GPT-3 (OpenAI, 2020), Google’s estimates (Patterson et al., 2021) give ~552 tCO₂e with a typical American energy mix. The same training on Google’s infrastructure would have produced ~21 tCO₂e according to these same authors — a factor of 26 that illustrates how much the datacenter’s location and electricity source change everything.
In short: training a model is the athlete’s Olympic effort (a single big peak that counts). Serving the model to users is the butcher’s daily routine (each transaction is small, but they add up day and night for years). After a year and a half, the butcher’s daily routine weighs as much as the initial training. And no one measures it systematically.
Meta published that training Llama 3 (405 billion parameters) required 6.4 million GPU-hours on H100s. At the typical consumption of an H100 (~700 W), this represents approximately 4.5 GWh — which, depending on the energy mix, amounts to roughly 2,000 to 3,000 tCO₂e. For GPT-4, OpenAI has published no data.
Why cooling has become a physical problem
A traditional datacenter consumes between 4 and 10 kilowatts per server rack. A GPU cluster for training large models reaches 100 to 130 kW per rack — a multiplication by ten to thirty. At these densities, air cooling is physically impossible: the required airflow rates cannot be achieved in an enclosed building.
The adopted solution is liquid cooling: plates in direct contact with the processors, through which water or a heat-transfer fluid circulates, dissipate heat far more effectively than air. In 2024, these systems represent 46% of the datacenter cooling market and have become the standard for all new hyperscale deployments. Microsoft, Google, and Meta have switched their AI clusters to liquid cooling. CoreWeave — a GPU cloud specialist — has operated all its new datacenters with liquid cooling since 2025, for NVIDIA GB200 NVL72 racks dissipating 130 kW.
To measure a datacenter’s energy efficiency, the historical indicator is PUE (Power Usage Effectiveness): the ratio between total energy consumed by the facility and energy consumed by IT equipment. A PUE of 1.0 is theoretically perfect — all energy goes to compute. A PUE of 2.0 means as much energy is dissipated in cooling as in useful compute. In 2024, the global average for traditional datacenters sits at 1.56. Modern hyperscalers reach 1.09 to 1.15 (Google reports a mean PUE of 1.09 across its global fleet).
But this metric has its limits: a datacenter with an excellent PUE running at low load can be less efficient than a less-optimized datacenter running at full power. Alternative metrics (DPPE, CUE, WUE) exist but have not yet established themselves.
The nuclear turn of hyperscalers
The constraint does not come from technology alone. Grid connection availability has become the main bottleneck: in Northern Virginia — the world’s datacenter hub — grid connection delays reach 5 to 7 years. In Europe, waiting queues in the FLAP-D hubs (Frankfurt, London, Amsterdam, Paris, Dublin) exceed 7 to 10 years. Amsterdam and Dublin have suspended new authorizations.
Faced with this constraint, hyperscalers have adopted a direct long-term supply strategy via PPAs (Power Purchase Agreements): electricity purchase contracts over 15 to 20 years signed directly with a producer, bypassing the grid. This mechanism explains the spectacular commitments of recent years.
In September 2024, Microsoft signs a 20-year PPA with Constellation Energy for the restart of the Three Mile Island plant — the same plant whose permanent shutdown had been decided in 2019 for economic reasons. The investment is 1.6 billion dollars for 835 MW of capacity, with a planned commissioning in 2028. This is the first restart of a fully shut-down nuclear power plant in the United States.
SMRs (Small Modular Reactors) are also attracting investment. In October 2024, Google signs the world’s first corporate SMR capacity purchase agreement with Kairos Power: 500 MW via 6 to 7 molten salt reactors, with the first reactor planned for 2030. Amazon contracts with X-energy for more than 5 GW of SMRs to be deployed by 2039. Oracle, for its part, is planning a triple-SMR complex to power a 1 GW datacenter as part of the Stargate project.
In total, major technology companies signed more than 10 GW of potential nuclear capacity in the United States in 2024–2025.
Water footprint: the overlooked externality
Most public discussion focuses on carbon, but datacenters also consume water — a great deal of it. Evaporative cooling through cooling towers is still used in many facilities, particularly in regions where water is less scarce.
At the scale of an individual model, Li et al. (2023) estimate that training GPT-3 required approximately 700,000 liters of water. More tangible at the individual scale: a conversation of 20 to 50 questions with a model of this size would consume approximately 500 ml of water — the equivalent of a plastic bottle. These figures are estimates and depend heavily on the type of cooling and electricity source.
Microsoft officially acknowledged a 34% increase in its water consumption between 2021 and 2022, directly tied to the training of GPT-4. The Patterns article (Cell Press, 2025) estimates the water footprint of AI at between 312 and 764 billion liters in 2025. Texas datacenters project up to 399 billion gallons of water by 2030. Operators frequently impose confidentiality agreements on local authorities, making public assessment of these consumptions difficult.
Unresolved controversies
The efficiency paradox
Do efficiency gains reduce total consumption? The answer, documented by academic research, is no. This phenomenon is known as the Jevons paradox: when a resource becomes cheaper to use, demand increases until it exceeds the savings achieved.
A recent example illustrates this: DeepSeek R1, a model less costly at inference, triggered an increase in inference demand, canceling out the energy savings per request. ACM SIGARCH notes that “efficiency alone will not solve the carbon problem of datacenters” without demand-limiting mechanisms (Souza et al., FAccT 2025).
The argument of efficiency per service rendered — one LLM request consumes between 0.001 and 0.01 kWh — deserves to be taken seriously and challenged simultaneously. Luccioni (2023) estimates that ChatGPT consumes approximately 10 times more energy per request than a Google search. Above all, the argument ignores the rebound effect: if each request costs less, usage volume explodes. Generative AI does not replace other activities — it creates new ones.
Carbon accounting in question
Hyperscaler claims of “100% renewable” rest on RECs (Renewable Energy Certificates) — certificates purchased separately from actual consumption, with no required temporal or geographic correspondence. A datacenter that buys RECs produced in summer in Arizona can use them to offset consumption in winter in Ohio.
Google unilaterally adopted a more demanding standard in 2020, hourly carbon matching (24/7 Carbon-Free Energy): carbon-free electricity must match actual consumption hour by hour. In 2024, the real energy mix of American datacenters remains dominated by natural gas (more than 40%), followed by renewables (24%), nuclear (approximately 20%), and coal (approximately 15%).
CSR reports confirm the trend: Google shows a 48% rise in emissions between 2019 and 2023 despite its climate commitments, and Microsoft +30% between 2020 and 2023. Offsets do not erase actual growth.
Ligozat et al. (2022) systematically analyzed these arguments in the literature and show that most environmental claims in AI rest on offsets that do not guarantee temporal and geographic causality.
Scope 3: what we don’t count
GPU manufacturing is rarely integrated into laboratory carbon budgets. An Nvidia H100 generates approximately 150 kg of CO₂e at production. A modern datacenter contains tens of thousands of them. Network infrastructure, undersea cables, equipment replacement cycles — all these items do not appear in the reported figures.
SMRs: when?
No commercial SMR is operational in the United States in 2025. The technology is in the regulatory certification phase at the NRC (Nuclear Regulatory Commission). The announced timelines — 2030 for the first Google-Kairos reactors, 2039 for the majority of Amazon’s capacity — raise a practical question: what will the actual energy mix of datacenters be during the 10 to 15 years of transition? The likely answer is: more natural gas, which increases net emissions in the short term.
What trade-off depending on hyperscaler context
| Context | Observed recommendation | Why |
|---|---|---|
| New datacenter, site available, < 12 months | Natural gas + cogeneration | Grid connection delay often exceeds the product horizon. Gas fills the gap. |
| Site available, grid capacity, 2-5 years | Solar/wind PPA + storage | Mature, financing available, but intermittency to compensate. |
| Growth > 1 GW, 5-10 years | Nuclear PPA (Three Mile Island, Susquehanna) or plant restart | Constant 24/7 load = perfect match for baseload nuclear. AWS, Microsoft, Meta have signed. |
| 2030-2035 horizon, massive capacity | SMR (small modular reactors, NuScale, X-Energy) | Promise of capacity in 50-300 MW increments. No commissioning confirmed before 2030. |
| Water constraint (arid regions, conflicts with agriculture) | Direct liquid cooling + dry adiabatic | Reduces water consumption by 70-90% vs cooling towers, at the cost of additional electricity (5-10%). |
Key takeaways
- Global datacenter consumption reaches 415 TWh in 2024 and could double by 2030 according to the IEA, driven by AI workloads growing at 30% per year.
- The footprint breaks down into training (massive one-time cost) and inference (low per request, but comparable over a model’s production lifetime).
- GPU clusters impose densities of 100 to 130 kW per rack — ten to thirty times more than conventional servers — making liquid cooling unavoidable for new infrastructure.
- Water consumption is the overlooked externality: ~500 ml estimated per conversation with a GPT-3-sized model; a 34% increase confirmed by Microsoft between 2021 and 2022.
- Hyperscalers are betting on nuclear (Three Mile Island restarted for Microsoft, SMR contracts at Google and Amazon) to free themselves from grid connection constraints.
- “100% renewable” declarations rest on certificate mechanisms decoupled from actual consumption; the effective mix remains predominantly fossil.
- Efficiency gains do not reduce total consumption — they stimulate demand (Jevons paradox), a conclusion academically documented and illustrated by the DeepSeek effect.