Dashboard

How Much Energy Does an AI Query Use? The Real Numbers

Google measured 0.24 watt-hours for a median text prompt. The same prompt is 0.10 Wh if you count only the chips, and that gap is the whole story.

Cecilia Iona
Cecilia Iona
Senior Editor, AI & Product
20 September 20261 min read

How much energy an AI query uses, on the best-sourced public number available, is 0.24 watt-hours for a median text prompt. That figure comes from Google, which in August 2025 published a measured breakdown for Gemini Apps, alongside 0.03 grams of CO2 equivalent and 0.26 millilitres of water per prompt. It is the most methodologically complete figure any provider has released.

It is also not the number most articles quote, and the gap between the two is the actually interesting part of this question.

How much energy an AI query uses depends on what you count

In the same disclosure, Google gives a second figure. If you count only the active power draw of the accelerators doing the inference, the TPU or GPU and nothing else, the same median prompt comes in at 0.10 Wh, 0.02 gCO2e, and 0.12 mL of water.

Same prompt. Same model. Less than half the energy. The only thing that changed is the accounting boundary.

What is counted

Energy

Emissions

Water

|---|---|---|---|

Accelerator only

0.10 Wh

0.02 gCO2e

0.12 mL

Full serving system

0.24 Wh

0.03 gCO2e

0.26 mL

Google's comprehensive figure adds four things the narrow one omits: the host CPU and RAM, idle machines held in reserve for failover and traffic spikes, data centre overhead such as cooling and power distribution, and the water used for that cooling. The company calls the narrow accounting "an optimistic scenario at best" that "substantially underestimates the real operational footprint of AI."

That is a provider volunteering that the flattering version of its own number is misleading, which is worth noting when you see any single figure quoted without its methodology.

What 0.24 watt-hours actually feels like

Small. A 2,000-watt kettle boiling for thirty seconds uses roughly 17 Wh, about seventy prompts' worth. An hour of a 100-watt television is around 417 Wh.

The honest framing is that per-query energy is close to irrelevant for an individual, and the aggregate is what matters at the scale providers operate. A million prompts a day at the comprehensive figure is 240 kWh per day, which is real infrastructure load. Your personal usage is not the story. Total demand is.

There is also a direction of travel. Google reports the median Gemini Apps text prompt dropped 33x in energy and 44x in total carbon footprint over twelve months, comparing May 2024 with May 2025, while response quality improved. Efficiency work compounds fast at this stage, which is part of why any energy figure older than about a year should be treated as historical rather than current.

Why your own app's number will be different

The 0.24 Wh figure is a median for short consumer text prompts on Google's infrastructure. Almost nothing about your workload matches that.

  • Output length dominates. Generation is sequential and each output token costs a full forward pass, which is why output tokens are priced higher than input tokens. A 2,000-token answer is not twice a 1,000-token answer in user-visible latency, but it is roughly twice the inference work.

  • Model size dominates too. A frontier model and a small one differ by more than an order of magnitude per token, for the same reason larger models carry a higher price per token.

  • Reasoning modes multiply everything. A model that generates hidden intermediate tokens before answering can do several times the work of a direct answer for the same visible output.

  • Retrieval and tool calls add rounds. Each one is another pass through inference, plus whatever the tool itself costs.

This is the practical reason a small model is often the environmentally and financially better answer when it is good enough for the task. Small language models are not a compromise everywhere, and the cases where they are sufficient are more common than most builders assume.

Using this honestly in your own claims

If a client or a tender asks about the footprint of AI in your product, here is what holds up and what does not.

Defensible: citing a provider's published methodology and figure, stating the date, and noting which accounting boundary it uses. Estimating your own volume by token counts and applying a published per-token or per-prompt figure with the caveat that it is an approximation.

Not defensible: quoting a single headline number with no methodology, comparing figures from two providers that used different boundaries, or using a 2023 estimate as though it describes current systems. Early public estimates were frequently an order of magnitude off in both directions, and some of the most-shared numbers were extrapolations from training costs rather than inference.

If cost is the real question behind the environmental one, and it often is, the levers that reduce energy also reduce your bill. Shorter outputs, smaller models where they suffice, and caching are the same three moves in both cases, which is covered in more depth in reducing AI API costs.

Frequently asked questions

How much energy does one ChatGPT query use?

OpenAI has not published a measured figure with a documented methodology comparable to Google's. Numbers circulating for ChatGPT are estimates by third parties, usually derived from assumed hardware and traffic. Treat them as informed guesses, not measurements.

Older comparisons put an AI query at roughly ten times a search. At 0.24 Wh that multiple no longer holds against commonly cited search figures, and the comparison was always shaky because the two do different amounts of work. The honest answer is that the gap narrowed substantially and the current numbers are close enough that the comparison stopped being useful.

Does running a model locally use less energy?

Usually more per query, not less. Consumer hardware is far less efficient per token than data centre accelerators running at high utilisation, and an idle local GPU still draws power. Local models win on privacy and offline access, not on energy.

What about the water figure?

The 0.26 mL is cooling water attributable to the prompt, roughly five drops. As with energy, it only becomes a meaningful quantity in aggregate, and it varies enormously by data centre location and cooling design.

Why do published estimates vary so much?

Three reasons: different accounting boundaries, different models and prompt lengths, and different vintages. Any two figures that differ by an order of magnitude are usually measuring different things rather than disagreeing about the same thing.

How did this land?

About the author

Cecilia Iona
Cecilia Iona

Senior Editor, AI & Product

Cecilia leads the Swarmz editorial desk. She has spent a decade turning complex AI and product topics into writing people actually finish, and she owns the blog's quality bar.

Share

Get the next post in your inbox

One email a month. Product updates, engineering posts, and the best of Built with Swarmz.

I agree to receive emails about AI building tips and Swarmz product news. Unsubscribe any time.