In brief

Every time a large language model is trained, tens to thousands of tonnes of CO₂ are emitted. Every query sent to a model like ChatGPT consumes water — a few centilitres per exchange, accumulated over billions of calls. This topic has been covered in the media since 2019, but the figures often circulate out of context. This article disentangles the real orders of magnitude, the industry arguments, and the questions that research has not yet resolved.


Training vs inference: two types of costs

In plain terms : training a model is like building a factory — a massive, one-off, noisy construction site. Using it is like running the factory every day — discreet on a per-unit basis, but cumulated over millions of products, it quickly exceeds the initial construction cost.

Understanding the environmental footprint of AI starts with distinguishing two radically different phases.

Training is the one-time but massive cost of training a model. It mobilizes thousands of GPUs for weeks or months. This is the spectacular figure cited by press articles.

Inference is the cost of each user query. Low on a per-unit basis, but accumulated over billions of daily calls, it represents a growing share — sometimes comparable — of a model’s total footprint over its lifetime.

A landmark 2022 paper by Luccioni et al. on BLOOM, a 176-billion-parameter model trained on the Jean Zay supercomputer in France, shows that eighteen months of inference produced a carbon footprint comparable to training itself (24.7 tCO₂e for training). It is the first academic study to document this balance on a model of this size.


Orders of magnitude for training

In plain terms : training GPT-3 in a typical US datacenter emits as much CO₂ as 120 combustion-engine cars each driving 15,000 km for a year. The same training run at Google, with a cleaner electricity mix, drops to 5 cars. Same computation, 26 times fewer emissions — just because the electricity does not have the same color.

In 2019, Strubell et al. (Carnegie Mellon) published a measurement that would make an impression: training an NLP model with neural architecture search can emit up to 284 tonnes of CO₂, roughly equivalent to five round-trip transatlantic car journeys over their full lifetime. This figure corresponds to an extreme case — neural architecture search (NAS), far more expensive than standard training. It nevertheless had the merit of making visible a cost that academic publications had never habitually reported.

For GPT-3 (OpenAI, 2020), Google’s estimates (Patterson et al., 2021) give ~552 tCO₂e with a typical US energy mix. The same training on Google infrastructure, partially fed by renewables and carbon-offset, would have produced ~21 tCO₂e according to the same authors. A factor of 26 — illustrating how much the datacenter location and electricity source change everything.

For more recent models, official data is scarce. Meta published that training Llama 3 (405 billion parameters) mobilized 6.4 million GPU-hours on H100s. At a typical H100 power draw (~700 W), this represents approximately 4.5 GWh — which, depending on the energy mix, amounts to roughly 2,000 to 3,000 tCO₂e. For GPT-4, OpenAI has published no data. These figures remain third-party estimates.


Water consumption: the forgotten figure

In plain terms : around twenty questions asked to ChatGPT “drink” the equivalent of one plastic water bottle. Multiply that by billions of queries per day: we are talking about Olympic swimming pools evaporated every hour to cool the servers.

Water is the poor relation of the environmental debate on AI. Water-cooled datacenters consume fresh water — both directly on-site (cooling towers) and indirectly via the power plants that supply them.

Li et al. (2023), in the first systematic academic study on the subject, estimate that training GPT-3 required approximately 700,000 litres of water. More tangible at the individual scale: a conversation of 20 to 50 questions with a model of this size would consume approximately 500 ml of water — the equivalent of a plastic bottle. These figures are estimates and depend heavily on cooling type and electricity source.

Microsoft officially acknowledged a 34% increase in its water consumption between 2021 and 2022, directly linked to training GPT-4. Goldman Sachs (2024) projects that generative AI datacenters could consume between 15 and 23 billion litres of water per year by 2027.


Per-request efficiency: solid argument or smokescreen?

In plain terms : each query consumes less, but we run many more of them. It is the Jevons paradox of the efficient engine: the car consumes 30% less per kilometer, but we drive 50% more — in total, we consume more.

The most common argument in industry communications is that of efficiency per unit of service delivered. A query to an LLM consumes between 0.001 and 0.01 kWh — less, according to some estimates, than a standard Google search.

This argument deserves to be taken seriously, and challenged simultaneously.

What is solid: per-request efficiency of models is improving rapidly. Models like Llama 3 8B or Mistral 7B achieve performance comparable to GPT-3.5 on many tasks with 10 to 100 times less energy. Algorithmic optimization (Chinchilla law, 2022) refocused research on compute-optimal training rather than the parameter count race.

What raises questions: Luccioni (2023) estimates that ChatGPT consumes approximately 10 times more energy per query than a Google search — contrary to what some corporate communications claim. Above all, the unit efficiency argument ignores the rebound effect: if each query costs less, usage volumes explode. Generative AI does not replace other activities — it creates new ones. Total consumption grows even as unit consumption falls.

The International Energy Agency (IEA, 2024) projects that global electricity consumption by datacenters could double between 2022 and 2026, from 240 TWh to between 500 and 1,000 TWh, with AI as the primary driver of this growth.


Carbon offsets: the question of rigour

In plain terms : buying a renewable energy certificate is like paying someone cycling on the other side of the country to offset your own diesel car trip. The accounting balance is neutral. The trip itself, however, did take place on diesel.

Google, Microsoft, and Meta purchase renewable energy certificates (RECs) to “neutralize” their emissions. Academic research raises a structural limitation of this practice: buying a wind energy certificate produced in Texas does not mean that the datacenter in Virginia is actually running on wind power at the moment the query is processed. Ligozat et al. (2022) systematically analysed these arguments in the literature and show that most “greenwashing claims” in AI rely on offsets that do not guarantee temporal and geographic causality.

The more rigorous approach, known as “hourly matching” — hour-by-hour correspondence between local renewable production and consumption — is much rarer. Google partially practices it.

Furthermore, Google’s CSR reports show a 48% rise in its emissions between 2019 and 2023 despite climate commitments, and Microsoft shows +30% between 2020 and 2023. Offsets do not erase real growth.


Scope 3: what we don’t count

In plain terms : AI carbon accounting only counts the electricity consumed during operation. Manufacturing the chips, digging the tunnels for submarine cables, replacing hardware every three years — all of this stays off the books. As if we measured a car’s footprint without counting its construction.

GPU manufacturing is rarely integrated into laboratory carbon footprints. An Nvidia H100 generates approximately 150 kg of CO₂e in production. A modern datacenter counts tens of thousands of them. Network infrastructure, submarine cables, equipment replacement cycles — all are line items that do not appear in reported figures.


Where to focus reduction efforts depending on context

Three concrete levers, depending on the usage profile and available level of control:

ContextRecommendationWhy
Deployment of a proprietary LLM with high inference volumeFavor compact models (7-8B) on targeted tasksCumulative inference often exceeds training over 18 months (BLOOM, Luccioni 2022). A 10× smaller model is enough for 80% of cases.
Training a new model from scratchChoose the datacenter based on local energy mixFactor 26 between typical US mix and Google renewable (Patterson 2021). Location weighs more than model size.
CSR communication on carbon offsetsRequire “hourly matching” rather than RECsRenewable energy certificates guarantee neither temporal nor geographic causality (Ligozat 2022). Only hourly matching is verifiable.

Key takeaways

  • The environmental footprint of AI breaks down into two phases: training (one-time, massive cost) and inference (low unit cost, but accumulated over billions of requests).
  • Training a large model emits on the order of a few hundred to a few thousand tonnes of CO₂e — a figure that depends as much on datacenter location and electricity mix as on model size.
  • Water consumption is an often-ignored impact: ~500 ml estimated per conversation of 20–50 questions with a model the size of GPT-3.
  • The per-request efficiency argument is real but insufficient: unit gains are offset by usage growth (rebound effect), and total consumption increases.
  • Transparency is the central problem: virtually all data on proprietary models comes from the industry itself or from third-party estimates. No reporting standard is mandated.