AI Energy & Sustainability 2026: How Much Power Do Claude, ChatGPT, Gemini and Mistral Really Use?
Every prompt you send to Claude, ChatGPT, Gemini or Mistral runs on a GPU or TPU somewhere in a data centre — and in 2026 those data centres have become one of the fastest-growing electricity consumers on the planet. Anthropic and OpenAI now plan compute in gigawatts, the IEA expects data-centre demand to almost double by 2030, and the EU has started to label data centres like fridges. At the same time, the energy of a single prompt has fallen dramatically. Here is what the numbers really say — and what a company can do to use AI responsibly without giving up its benefits.
The big picture: AI is now an electricity story
The International Energy Agency's April 2026 update gives the clearest numbers. Data centres worldwide consumed about 485 TWh in 2025 — a 17% jump in a single year and more than five times the growth rate of total electricity demand. Facilities built mainly for AI used about 155 TWh, up roughly 50% in a year. In the IEA base case, total data-centre consumption reaches about 950 TWh by 2030, and AI-focused facilities roughly triple.
For scale: 950 TWh is more than the annual electricity consumption of Japan. Growth is concentrated in two countries — the IEA expects the United States to go from about 224 to 426 TWh and China from 117 to 277 TWh between 2025 and 2030. Europe grows more slowly, held back by grid connections and permitting rather than demand.
The gigawatt race: Anthropic, OpenAI and the build-out
The model labs no longer talk about servers but about power plants' worth of capacity. Anthropic has assembled a compute pipeline of more than 15 GW across about 15 US campuses, including agreements with Amazon for up to 5 GW and with Google and Broadcom for another 5 GW of TPU capacity; its Amazon campus in New Carlisle, Indiana is already among the largest AI data centres in the world at roughly 900 MW. OpenAI's Stargate commitments have passed their original 10 GW target, with further multi-gigawatt sites announced in Georgia and Ohio. Because gigawatt campuses take years to connect to the grid, both companies are also renting smaller 20–30 MW facilities to meet demand right now.
For businesses this matters in two ways: compute supply — and therefore price and rate limits — is tied to how fast the grid can grow, and the energy mix of those new sites (gas, nuclear, renewables) determines the real carbon footprint of the AI you buy. If you want the sector-side view, see our guide to AI in energy and utilities.
What does one prompt really cost?
Per-prompt numbers became public only in 2025, and they are much smaller than the viral claims of "a bottle of water per email" suggested:
| Provider | What is disclosed | Key figures |
|---|---|---|
| Full methodology for Gemini Apps (Aug 2025) | Median text prompt: 0.24 Wh, 0.03 g CO₂e, 0.26 mL water (~5 drops); 33× less energy than 12 months earlier | |
| OpenAI | CEO statement (Jun 2025) | Average ChatGPT query: ~0.34 Wh and ~0.32 mL water |
| Mistral | Peer-reviewed lifecycle analysis of Mistral Large 2 with ADEME and Carbone 4 | 400-token answer: 1.14 g CO₂e, ~45 mL water (incl. share of training and hardware); training + 18 months of use: 20.4 kt CO₂e, 281,000 m³ water |
| Anthropic | No per-prompt figure published | Independent researchers estimate under 1 Wh for a short Claude Sonnet query; 15+ GW compute pipeline |
| Manus & agents | No disclosure | An agent task chains dozens of model calls — footprint scales with the number of steps |
These figures are not directly comparable: Google reports a median, OpenAI an average, and Mistral a full lifecycle number that includes training and hardware manufacturing — which is why its CO₂ and water figures are higher. Mistral deserves credit as the first major lab to publish a peer-reviewed analysis of this kind.
The catch is in the right-hand side of the AI bars. Reasoning models that "think" for thousands of tokens, prompts that load a 200-page contract and autonomous agents that run for twenty minutes use 10 to 70 times more energy than a simple chat answer. As AI shifts from questions to long-running agents, the energy per task grows even as the energy per token falls.
Water and carbon: location matters more than the model
Data centres use water mainly for cooling, and the carbon footprint of a kilowatt-hour depends on the grid. In the EU, grid intensity ranges from under 50 g CO₂ per kWh in countries with mostly hydro, nuclear or wind to several hundred grams where coal and gas dominate. The same prompt can therefore have a ten-times different carbon footprint depending on the region it runs in. Choosing an EU region with a clean grid — something that also helps with data residency — is one of the simplest levers a company has.
The regulators are watching
- EU AI Act: providers of general-purpose AI models must document the known or estimated energy consumption of their models in their technical documentation. More in our EU AI Act compliance guide.
- Energy Efficiency Directive: data centres with more than 500 kW of IT power must report energy, water and renewable-share data to an EU database.
- EU data-centre sustainability label: in July 2026 the Commission published a revised draft of a mandatory rating scheme and electronic label for data centres; the first labels are expected in 2027.
- Sustainability reporting: companies that report emissions will increasingly need to account for cloud and AI services in their Scope 3 figures — ask your providers for data now.
Green AI is cheap AI: 7 levers that cut energy and cost
The good news: almost everything that lowers the footprint of AI also lowers the bill, because providers price by compute. The levers below are the same ones we use to optimise inference costs:
- Right-size the model. Use small, fast models (Claude Haiku, GPT mini models, Mistral Small) for classification, extraction and routing, and reserve frontier models for hard problems. Model routing has cut costs by over 85% in research settings while keeping ~95% of quality.
- Cache what repeats. System prompts, manuals and contracts sent again and again should use prompt caching — up to 90% cheaper on cached input.
- Batch what can wait. Overnight reports and document runs via a Batch API cost 50% less and let providers schedule work when capacity — and often clean power — is available.
- Control reasoning and output length. Set thinking budgets, ask for concise answers and stop agents after a sensible number of steps.
- Run small models locally where it fits. Small language models and quantisation (up to ~75% less memory at 4-bit) make on-premise or edge AI practical for simple tasks.
- Pick clean regions and transparent providers. Prefer EU regions with low grid intensity and providers that publish energy, water and emissions data.
- Measure. Track tokens, model mix and cost per task in a dashboard; convert to kWh and CO₂ with your provider's data for your sustainability report.
The bottom line
AI's energy problem is real but often misunderstood. A single prompt is cheap in energy — and getting cheaper each year — while the total demand of AI is rising fast and reshaping power grids. For a company, the responsible answer is not to avoid AI but to use it deliberately: the right model for each task, caching and batching by default, clean regions and measurement. Done well, green AI is simply efficient AI — better for the climate and for your margins.
Want efficient, measurable and sustainable AI in your company?
We help companies in Croatia and the DACH region design AI systems that are fast, affordable and responsible: model selection and routing, caching and batch pipelines, EU-region deployments and usage dashboards that feed your sustainability reporting. As a Claude Certified Architect based in Zagreb, we build with Claude and other enterprise models every day.
Talk to an AI consultant