How much energy does an AI query use? What we can and can't know
An honest guide to AI energy use per query — what providers disclose, what researchers estimate, what remains unknown, and how EcoRouter models it.
"How much energy does one AI question use?" is a fair question with an unsatisfying answer: it depends, and for your individual request, nobody outside the provider can measure it. What we do have is a small set of provider disclosures, independent estimates and research methods that let us reason about it. This guide separates what is measured, what is modeled and what is still unknown.
The short answer
Recent public figures for a typical text prompt cluster around a few tenths of a watt-hour:
- Google reported that the median Gemini Apps text prompt used about 0.24 Wh, measured in production and including idle capacity and data center overhead1.
- OpenAI's chief executive stated that an average ChatGPT query uses about 0.34 Wh, without publishing a method2.
- Epoch AI estimated roughly 0.3 Wh for a typical ChatGPT query using GPT-4o, working from first principles3.
- A study by Oviedo and colleagues, published in Joule in 2026, estimated a median of 0.31 Wh per query for frontier-scale models (over 200 billion parameters) on H100 nodes, assuming optimized large-scale serving, with an interquartile range of about 0.16 to 0.60 Wh4.
Older and widely repeated estimates were roughly ten times higher, around 3 Wh per request, derived from industry cost estimates rather than measurements5. The gap mostly reflects different assumptions about hardware, power draw, answer length and model size, not a contradiction in the physics34.
None of these numbers describes your query. Google's is a measured median for one product on its own infrastructure; OpenAI's is an average with no published method; Epoch AI's and Oviedo's are modeled estimates. A long prompt, a long answer or a reasoning model that "thinks" through many hidden tokens can use several times more.
Why a single request can't be measured
Measuring the electricity of one request would mean reading the meter on the hardware that served it, at the moment it ran, and attributing a fair share of everything around it. In practice that is very hard:
- Requests are batched. One accelerator processes many users' requests together, so its power draw is shared.
- Hardware is shared and varies. The same model may run on different chips in different data centers, at different utilization levels.
- Overhead is real. Cooling, power conversion and idle capacity held ready for demand spikes all consume energy that no single request "owns."
- Providers don't expose it. None of the providers EcoRouter routes to returns an energy figure with a response, and none publishes the sizes of these models.
So any per-request energy figure from a third party is modeled, not measured, however precise it looks.
Measured, modeled and unknown
It helps to sort every number into one of three categories:
- Measured: observed directly. For an API request, that includes the input and output token counts the provider reports and the time the request took.
- Modeled: calculated from measured values plus stated assumptions — for example, energy estimated from token counts and an assumed model size and hardware profile. Only as good as its assumptions, which should be published.
- Unknown: the actual model size, which hardware ran the request, its utilization at that moment, and the carbon intensity of the grid supplying it.
A credible energy claim tells you which category it belongs to. A figure that blurs them — a modeled estimate presented to three decimal places with no stated method, for example — deserves skepticism.
What drives the energy of a query
Several factors move the number:
- Model size. Computation per token scales with the parameters a model activates, at about two floating-point operations per active parameter per token36. A model with twice the active parameters does roughly twice the work per token.
- Answer length. Input is processed in parallel, but output is generated one token at a time, and this decoding stage often dominates energy use4. Reasoning models that generate many intermediate tokens can multiply energy per query: one study estimated that queries 15 times longer than typical raise median energy about 13-fold4.
- Prompt length. Long documents and conversation history add input tokens, and very long contexts cost more than linearly3.
- Hardware and utilization. Newer accelerators and well-batched serving do more work per watt-hour.
- Data center overhead. Cooling and power delivery add to every request.
- Task type. Research comparing many models found generative tasks considerably more energy-intensive than classification-style tasks, and image generation more intensive than text7.
Small per query, large in aggregate
A few tenths of a watt-hour is small next to everyday appliances. The concern is scale: billions of queries a day, on top of training and growing AI workloads. In its 2026 update, the International Energy Agency estimated that data centers used about 485 TWh of electricity in 2025, and projected that this could roughly double to about 950 TWh by 2030, around 3% of global electricity demand, with electricity use by AI-focused data centers growing much faster and roughly tripling over the same period8. Those are totals for all data center work, including training, inference and non-AI services, not a per-query figure.
That is why per-query efficiency matters even when each query is small: the cheapest unit of energy to manage is the one a request never needed.
How EcoRouter estimates the energy of an answer
EcoRouter does not measure the electricity of any request, and it does not estimate carbon or water. What its Eco Receipt shows is a modeled, low-confidence energy estimate in watt-hours, published in full in the Energy Consumption & Savings methodology.
The method follows Epoch AI's first-principles approach3. For each answer, EcoRouter takes the measured input and output token counts and estimates the energy needed to process them, using an assumed number of active parameters for the model that answered, a reference accelerator's throughput and utilization, and an average power figure per accelerator that includes server and data center overhead. It runs the same calculation for a configured frontier baseline model on the same tokens. The difference is the estimated energy saved.
Because no provider EcoRouter routes to publishes model sizes, each model is placed in one of four size classes — compact, mid-size, large and very large — based on how the provider positions it. That judgment is the largest uncertainty in the method, and it is labeled as such.
EcoRouter also refuses to claim a saving that its own uncertainty could erase. Each estimate is allowed to range from half to double its central value, and a saving is shown only if it survives the worst case: the baseline at half its estimate must still exceed EcoRouter's route at double. Otherwise the receipt says the saving is uncertain.
The estimate covers running the answering model, including any hidden reasoning tokens billed as output, plus data center overhead through the power assumption. It excludes model training, hardware manufacturing, EcoRouter's own servers, network transfer, your device, and the energy of a live web search engine. Reused answers are handled separately: when EcoRouter serves a saved answer, no model runs, and that saving is not yet estimated.
Takeaway: how to read any AI energy claim
Before trusting an energy figure for AI — ours included — check four things:
- Is it measured or modeled? If modeled, are the assumptions published?
- What is in scope? Inference only, or training and hardware too? Overhead included?
- Is it a median, an average or one request? These differ widely.
- Does it show uncertainty? A low-confidence estimate presented with false precision is a warning sign.
The most useful habit is the simplest: prefer sources that tell you what they don't know. The EcoRouter FAQ explains what we do and don't claim, and you can see a modeled estimate on a real answer by asking EcoRouter a question.
References
-
Google (Elsworth et al.), "Measuring the environmental impact of delivering AI at Google Scale" (2025). https://arxiv.org/abs/2508.15734 (opens in a new tab) ↩
-
Sam Altman (OpenAI CEO), "The Gentle Singularity", personal blog (10 June 2025). https://blog.samaltman.com/the-gentle-singularity (opens in a new tab) ↩
-
Epoch AI (Josh You), "How much energy does ChatGPT use?" (7 February 2025). https://epoch.ai/gradient-updates/how-much-energy-does-chatgpt-use (opens in a new tab) ↩ ↩2 ↩3 ↩4 ↩5
-
Joule (Oviedo et al.), "Energy use of AI inference, efficiency pathways, and test-time scaling" (2026), article 102430. https://doi.org/10.1016/j.joule.2026.102430 (opens in a new tab) (preprint: https://arxiv.org/abs/2509.20241 (opens in a new tab)) ↩ ↩2 ↩3 ↩4
-
Joule (Alex de Vries), "The growing energy footprint of artificial intelligence" (2023). https://doi.org/10.1016/j.joule.2023.09.004 (opens in a new tab) ↩
-
arXiv (Kaplan et al., OpenAI), "Scaling Laws for Neural Language Models" (2020). https://arxiv.org/abs/2001.08361 (opens in a new tab) ↩
-
ACM FAccT (Luccioni, Jernite and Strubell), "Power Hungry Processing: Watts Driving the Cost of AI Deployment?" (2024). https://arxiv.org/abs/2311.16863 (opens in a new tab) ↩
-
International Energy Agency, "Key Questions on Energy and AI" (2026 update to the 2025 "Energy and AI" report). https://www.iea.org/reports/key-questions-on-energy-and-ai (opens in a new tab) ↩