Public EcoRouter Knowledge · AI-generated content ·
How do u calculate energy savings / estimates per query?
The question
How do u calculate energy savings / estimates per query?
EcoRouter calculates energy estimates using a published modeling method (based on research by Epoch AI), rather than direct physical measurements, because AI providers do not report per-request electricity data.
Here is how the calculation works:
1. Estimating Energy Used
For each query, the energy is modeled based on the number of tokens processed and the model's estimated size:
- Token counts: Input tokens (reading the prompt) and output tokens (generating the response, including any hidden reasoning tokens) are recorded.
- FLOP calculation: Each token requires roughly per active parameter in the model's estimated size class (ranging from Compact at ~15B parameters to Very Large at ~200B).
- Hardware & facility power: The compute is modeled at NVIDIA H100 throughput levels, factoring in estimated hardware utilization (10% generating, 28% reading) and data-centre power overhead (around 1,500 W per accelerator).
2. Calculating Energy Saved
To estimate savings, the model that answered is compared against a frontier baseline:
- The baseline represents sending the query to a flagship general-purpose model (Large size class).
- The calculation assumes the baseline model would have read the exact same prompt and produced a response of the same length.
- The difference between the baseline's estimated energy and the actual model's estimated energy gives the potential saving.
3. Accounting for Uncertainty
Because model architectures and real-time data-centre conditions are proprietary, these figures are low-confidence estimates with an assumed uncertainty range of 0.5× to 2× the central value.
To avoid overstating savings:
- A saving is only reported as "saved" if the baseline's lower bound is strictly greater than the chosen model's upper bound (
baseline × 0.5 > answer × 2). - If the uncertainty ranges overlap, the result is marked as "uncertain" rather than claiming a saving.
What Is Not Included
The estimate covers the inference compute and data-centre overhead for the query itself. It does not include model training, hardware manufacturing, network transmission, user devices, search engine energy, carbon emissions, or water use.
You can read the full documentation, parameters, and sources in the EcoRouter Energy Methodology.
Answering this used an estimated 75% less than EcoRouter’s configured frontier baseline
- Routed to
- Balanced
- Tokens
- 1,850 in / 962 out
- Cost
- $0.0050
- Baseline
- $0.02
Figures recorded by EcoRouter when this answer was generated, and fixed at that moment. Cost comparisons are against a configured reference model, not a measurement of electricity, carbon or water.
Ask about this
Ask EcoRouter a follow-up using this Knowledge as context. Nothing becomes public unless you choose to publish it.
Ask a follow-upComments
No comments yet.