Abstract
Large-language-model inference is now a first-order energy and cost question for digital organizations [1][2]. The default response — asking a frontier model the same organizational question many times — wastes both tokens and watts. This working paper proposes a different architecture. A Human-Enhanced Agent (HEA) compiles an organization’s knowledge, voice, and governed skills once, then answers many subsequent operations on smaller models. We call this compile once, serve many.
We model four client scales (105 to 108), treat demand as a mix of customer chat, email classification, reply drafts, and other backoffice AI, and price dollars as electricity rather than vendor token list prices. Carbon saving has two explicit parts: fewer kWh from model downshift, plus the additional intensity effect of scheduling eligible automation off-peak while live work keeps its natural hour mix. At one million clients, the mid case moves from a 3.15 ktCO2 same-workload frontier footprint to a 1.10 kt HEA footprint: 1.80 kt avoided through lower demand plus 0.26 kt through scheduling, or 2.05 kt combined. The future visitor-Q&A case avoids 31.99 kt against its 40.76 kt same-workload frontier counterfactual, but its 8.78 kt HEA footprint is still 5.62 kt above today’s smaller mid frontier footprint. The conversation-centric case avoids 38.00 kt against its 61.87 kt counterfactual while reporting about 139,000 legacy contacts per client separately from energy. At 100 million clients, the future and conversation-centric cases avoid about 3.20 and 3.80 MtCO2/year; their 10-year Net Present CO2 Savings at 7.6% are about 21.86 and 25.96 Mt. These are scenario counterfactuals, not production measurements or adoption forecasts. The 2030 Agenda is a policy frame, not a claim that HEAs achieve the Sustainable Development Goals.
Executive summary
- The claim that survives review. Organization-specific AI can run on mini- and nano-class models when knowledge is compiled once. Savings exist only for operations that would otherwise have hit a frontier model.
- Price electrons, not SKUs. Token list prices move with product marketing and include GPU capital and margin. This paper’s primary money unit is kWh × an industrial or datacenter tariff (default USD 0.10 / kWh).
- Two carbon mechanisms, kept separate. Smaller models reduce demand. Email sync, classification, Foundry compile, and queued jobs can then move out of the natural live hour mix. With 80% of live/frontier work on-peak and 100% execution of eligible scheduling, effective HEA off-peak share is about 40% in the mid case, 34% in future, and 27% in conversation-centric.
- Carbon is not automatic. Peak-gas / off-peak-baseload is one named scenario. On solar-day / gas-night grids, shifting to “off-peak” raises intensity. Cost spread and carbon spread are different facts.
- Scale. At 100 million clients, the mid, future, and conversation-centric cases avoid about 0.205, 3.20, and 3.80 MtCO2/year against their same-workload frontier counterfactuals. Ten-year NPCS at 7.6% is about 1.40, 21.86, and 25.96 Mt. Future demand still raises the absolute HEA footprint above today’s smaller workload; avoided future emissions are not the same as an absolute reduction from today.
- Conversation-centric is a service count. The future case asks what happens if website Q&A absorbs ChatGPT. A conversation-centric case asks what happens if that same conversation also absorbs email, forms, pages, and phone. Those legacy contacts are not converted to frontier watt-hours. Energy save can stay near the mid 57% while the organization still retires a large contact pile.
2030 Agenda — policy frame, not a scorecard
The United Nations 2030 Agenda for Sustainable Development is the intergovernmental frame this paper uses when it talks about energy, infrastructure, consumption, and climate [15]. Official goal titles are used below. This paper does not claim that Human-Enhanced Agents achieve, deliver, or are aligned-certified against those goals. It does not reproduce UN icons or artwork.
Goal 7, Affordable and clean energy. The energy tables ask whether compiled mini/nano inference uses fewer watt-hours than the frontier calls it replaces, and whether batch work can sit in cheaper hours. That is an efficiency question inside Goal 7, not a renewable-generation claim.
Goal 9, Industry, innovation and infrastructure. Compile-once, serve-many is an infrastructure choice: put organizational knowledge in a package, then serve many operations on a smaller model, instead of re-deriving the same context on a frontier assistant.
Goal 12, Responsible consumption and production. The conversation-centric case counts contacts that would otherwise have been email, a form, a webpage visit, or a phone call. One conversation can replace a loop of those channels. The paper reports the contact count. It does not invent a kWh for a mailbox server or a call centre.
Goal 13, Climate action. The carbon columns are Goal 13 material only under a named hour-mix. Off-peak is not automatically cleaner. Table 3 is a scenario, not an inventory line.
Goal 8, Decent work and economic growth — handled carefully. Conversation-centric service changes how customer work is done. This paper does not count jobs created or lost, and it does not treat labour displacement as an energy saving.
1. Introduction
Training energy for foundation models has received more scholarly attention than inference, yet inference now dominates the operational footprint of deployed AI [2][3]. Organizations increasingly route internal questions, mailbox triage, and customer conversations through frontier assistants. Each call re-derives context that the organization already owns.
HEA World’s product architecture was built for a different constraint: token throughput, not servers, is the first scalability ceiling of conversational AI. The practical response is not a larger quota. It is fewer tokens per useful operation, on a smaller model, with the organization’s knowledge already compiled.
This paper is a working paper. It does not claim a measured production energy audit. It states a substitution thesis, a transparent scenario model, and the conditions under which the thesis fails.
2. Related work
Strubell et al. established that training large NLP models has a material energy and carbon cost [4]. Schwartz et al. argued for “Green AI”: reporting efficiency alongside accuracy [5]. Patterson et al. and Luccioni et al. refined training-footprint estimates and showed how location and generation mix dominate the result [3][6].
Deployment is a different problem. Luccioni, Jernite, and Strubell show that inference energy varies more by task and model class than by a single “ChatGPT query” number [2]. Dodge et al. measure cloud-instance carbon intensity as a function of region and time [7]. The IEA’s Energy and AI assessment places datacenter and AI electricity in the hundreds of terawatt-hours by the late 2020s, with large uncertainty [1].
This paper adds an architectural lever that those studies treat only implicitly: moving intelligence from model weights into a compiled organizational package, then downshifting the live model. We also refuse to treat API dollars as joules. The 2030 Agenda is cited as the public policy vocabulary for energy, infrastructure, consumption, and climate — not as an evaluation rubric for the product [15].
3. Architecture and the substitution thesis
A Human-Enhanced Agent is an owned representative shaped by an organization’s knowledge, voice, and human oversight [8]. In the HEA World stack, two layers matter for energy:
- Foundry (compile). Facts, summaries, embeddings, and localization. Infrequent. Default routing is already a mini-tier model, not a frontier model per document.
- Runtime (serve). Live chat defaults to a mini-class model (currently
gpt-4.1-mini). Language identification, extraction, and confirmation use a nano-class model. The corpus is not resent; retrieval sends slices plus compiled personality.
Figure 1 separates the two substitutions this architecture is for. First, each later chat turn is cheaper because the organization is already compiled (mini answer, nano helpers, retrieved slices). Second, who runs the model changes: one shared package plus lean multi-user automation replaces N people each driving an intensive frontier assistant (Copilot- or ChatGPT-class). The first is a per-turn watt-hour save. The second is a seat-and-automation save. Both still count only the replaceable share.
Demand is not “chat turns.” A working day also includes mailbox sync and classification, reply drafts, and other backoffice AI. The mid case in this paper is 10 daily users × 15 operations × 220 workdays: 25% chat, 40% classify, 20% draft, 15% other backoffice. That mix is a conservative operator-day picture. It is not the HEA World business-plan chat funnel.
Savings are not created by counting HEAs. They are created by the replaceable share — operations that would otherwise have been frontier calls.
Today, many visitor widget turns are still induced: they would not have happened. That is why the mid case treats only 30% of chat as replaceable. The substitution thesis changes if the website HEA becomes the default place to ask an organization a question — the turn that today goes to ChatGPT. Then customer Q&A volume is visitor-driven, not headcount-driven, and the replaceable share of chat should rise. Employee drafts and Copilot-style internal questions remain a displacement pool. The model keeps both shares explicit.
3.1 Future demand: visitor Q&A and lean automation
HEA World’s internal business-plan visitor funnel converts monthly unique visitors into conversations and turns (engagement ratio × sessions × turns per session). An informational site at 10,000 monthly uniques and a 2.5% engagement ratio produces about 250 engaged visitor conversations and 2,500 turns per month. The mid-case staff chat slice is about 688 turns per month. If the HEA becomes the primary Q&A surface, engagement can look more like support or a chat-native profile. The named future case therefore adds 8,000 visitor turns per client per month (about 96,000 per year), on top of the unchanged 10 × 15 staff day, and treats 65% of chat turns as replacing a frontier assistant.
The same growth path adds lean, cron-shaped automation: extra mailbox classification, queued drafts, and other backoffice jobs (CRM sync, briefing, skill runs). Those operations can move off-peak; live visitor chat cannot. The future case therefore does two different things at once. Demand reduction versus the same-workload frontier counterfactual rises to about 30.6 ktCO2 at one million clients. The kWh-weighted schedulable share falls from about 25% in mid to 18%, because chat dominates energy, but absolute schedulable kWh still rises about fivefold. Moving all eligible automation off-peak adds about 1.39 kt, for 31.99 kt combined. “We will move more work off-peak” is true for eligible watts; it does not mean that live chat becomes schedulable.
3.2 Conversation-centric engagement and legacy service contacts
The future sentence above still understates a conversation-centric organization. Today’s customers already ask questions; they do it by email, web form, a hunt through webpages, and the phone. Those contacts are IT and human operations. They are not frontier-model watt-hours. Folding them into the ChatGPT column would invent energy that was never spent on a frontier assistant.
HEA World’s support-site funnel on 10,000 monthly uniques is engagement ratio 0.25 × 2.5 sessions × 4 turns, or 25,000 visitor turns per client per month (about 300,000 per year), on top of the same 10 × 15 staff day. The named conversation-centric case treats that turn pile as three explicit shares:
- 35% frontier-replaceable — the turn that would otherwise have been ChatGPT or a Copilot-class assistant. This share enters the energy tables.
- 45% legacy-replaceable — one conversation counted as one contact-equivalent that would otherwise have been email, a form, a webpage visit, or a phone call. The model does not split those four channels. This share is not converted to frontier watt-hours.
- 20% induced — the widget created the turn. No frontier energy and no legacy contact.
At one client that is about 25,700 chat turns per month, about 108,000 frontier-replaced turns per year, and about 139,000 legacy contact-equivalents per year. At one million clients the legacy column is about 139 billion. HEA still pays for every turn (about 76 kWh per client per year; 76 GWh at one million). Demand reduction avoids about 36.33 ktCO2 at one million clients and scheduling adds about 1.67 kt, for 38.00 kt combined against a 61.87 kt same-workload frontier footprint. The resulting HEA footprint is 23.87 kt — 20.71 kt above today’s smaller 3.15 kt mid frontier footprint. The schedulable share falls to about 9% because live chat dominates, even while absolute schedulable kWh rises slightly.
Conversation-centric engagement is a contact count. It is not a claim that email servers, forms, pages, and phone switches used the same watt-hours as a frontier assistant.
4. Methods
4.1 Energy
Published per-query watt-hours for frontier assistants span roughly 0.3–3 Wh [2][9]. We treat 1.5 Wh as the mid frontier query and scale other operations by a relative token-intensity coefficient (frontier = 1.0, mini = 0.25, nano = 0.08, mini-band draft = 0.22). These coefficients are conservative relative to list-price ratios (mini is about 5× cheaper than gpt-4.1; nano about 20×) because energy does not inherit vendor prices.
A typical HEA chat turn uses the platform’s v2 benchmark shape: 3,250 input and 450 output tokens on mini, plus bounded nano helper calls. Classification is mostly nano; drafts use a mini-band model with a longer grounded prompt.
4.2 Money
Primary dollars:
electricity $ = kWh × (peak share × P_peak + off-peak share × P_off)
The flat default is USD 0.10 / kWh, an illustrative industrial / datacenter figure in the Eurostat and EIA non-household band [10][11]. Time-of-use uses USD 0.18 peak and USD 0.06 off-peak unless the reader changes the calculator. Token / API dollars are reported only as a footnote. They are typically two orders of magnitude larger because they price GPUs and margin.
4.3 Carbon
tCO2 = kWh × (peak share × CI_peak + off-peak share × CI_off)
Default thesis intensities are 400 gCO2/kWh on-peak (modern gas CCGT, IPCC-order operational factor) and 80 gCO2/kWh off-peak (nuclear / hydro / wind mix) [12]. A “tomorrow” case uses 25 g off-peak. An inverted case (160 g day / 380 g night) represents solar-heavy days and gas-heavy nights observed on some European grids [13][14]. Factors are operational, not a full life-cycle assessment of GPUs.
The result is decomposed rather than hidden in one total:
demand reduction = (frontier kWh − HEA kWh) × frontier carbon intensity
scheduling effect = HEA kWh × (frontier carbon intensity − scheduled HEA carbon intensity)
combined avoided = demand reduction + scheduling effect
4.4 What can move off-peak
Email sync, classification, Foundry compile, and part of queued draft/backoffice work are schedulable. Live chat is not. The natural live hour mix follows the frontier baseline: 80% on-peak and 20% off-peak by default. The kWh-weighted schedulable share is about 25% in mid, 18% in future, and 9% in conversation-centric. The default schedules 100% of that eligible work off-peak. Effective HEA off-peak share is therefore natural off-peak + schedulable share × execution × (1 − natural off-peak): about 40%, 34%, and 27%, respectively. The calculator exposes the live hour mix and scheduling execution separately. If the grid is dirtier off-peak, the scheduling effect becomes negative while demand reduction remains visible.
4.5 Legacy contacts are not a second energy column
Where a band sets a legacyReplaceable share, those operations are counted as contact-equivalents and excluded from baseWh. HEA still incurs watt-hours for serving the conversation. The energy save percentage is therefore 1 − HEA Wh / frontier Wh on the replaceable slice only. A large legacy pile can coexist with a mid-like save percentage. Phone minutes, mailbox-server electricity, and page-view joules are out of scope. The four-channel mix (email / form / page / phone) is not estimated.
4.6 Horizon, cars, and Net Present CO2 Saving
Tables 1–5 are a yearly flow. Compile-once, serve-many is infrastructure: if the architecture remains valid and the replaceable mix does not collapse, that flow repeats. For a constant annual combined avoided emission S held for n years:
cumulative = n × S
NPCS = S × Σt=1…n (1 + r)−t = S × [1 − (1+r)−n] / r
The default discount rate is r = 7.6%, the UNEP Emissions Gap 1.5 °C annual-cut pathway used in Payen (2021) as Net Present CO2 Saving [16][17]. A later tonne is worth less than a tonne saved this year because delayed cuts must then be steeper. First-year save is discounted one step (end-of-year), the same convention that maps a 2030 bullet to a 10-year discount. This is not a financial NPV and not a social cost of carbon. Electricity dollars in the horizon table stay undiscounted kWh × tariff. Car equivalents use 2 tCO2/year per EU passenger car; a yearly row is car-years, not cars permanently removed. Installed base, mix, and grid intensity are held constant across the horizon. The recommended reading window is 10 years. A 25-year row is a sensitivity toward 2050, not a product-lifetime claim. The calculator also exposes r = 17% (a 2023 PwC update of the required annual cut, noted in Payen 2021/2023).
5. Results
Unless noted, tables use the mid intensity band, USD 0.10 / kWh flat, USD 0.18 / 0.06 time-of-use, live/frontier work 80% on-peak, 100% scheduling execution for eligible automation, and the peak-gas / off-peak-baseload carbon case. One HEA per client. Foundry amortization is included in HEA kWh. Frontier columns count only the replaceable slice.
5.1 Energy and flat electricity dollars
| Clients | HEA ops / year | HEA energy | Net energy | Flat net $ | Energy save |
|---|---|---|---|---|---|
| 100,000 | 3.3 billion | 404 MWh | 534 MWh | $53.4k | 57% |
| 1 million | 33 billion | 4.04 GWh | 5.34 GWh | $534.5k | 57% |
| 10 million | 330 billion | 40.4 GWh | 53.4 GWh | $5.34M | 57% |
| 100 million | 3.3 trillion | 404 GWh | 534 GWh | $53.4M | 57% |
Table 1. Mid case. Per client, HEA electricity is 4.04 kWh/year; net versus replaced frontier work is 5.35 kWh/year. Percentage save does not change with client count.
5.2 Time-of-use spread
| Clients | Flat net $ | TOU, natural hours | TOU + eligible scheduling | Scheduling $ from spread |
|---|---|---|---|---|
| 100,000 | $53.4k | $83.4k | $93.1k | $9.69k |
| 1 million | $534.5k | $834k | $931k | $96.9k |
| 10 million | $5.34M | $8.34M | $9.31M | $969k |
| 100 million | $53.4M | $83.4M | $93.1M | $9.69M |
Table 2. Mid case. kWh do not change between the columns. Natural hours put frontier and unscheduled HEA work on the same 80/20 hour mix. Eligible scheduling moves all schedulable HEA automation off-peak, producing a 40% effective HEA off-peak share.
5.3 Carbon dioxide
| Clients | Demand reduction | Scheduling | Combined avoided CO2 | EU-car years |
|---|---|---|---|---|
| 100,000 | 180 t | 25.8 t | 205 t | ~100 |
| 1 million | 1.80 kt | 258 t | 2.05 kt | ~1,000 |
| 10 million | 18.0 kt | 2.58 kt | 20.5 kt | ~10,000 |
| 100 million | 180 kt | 25.8 kt | 205 kt | ~103,000 |
Table 3. Mid case; peak 400 gCO2/kWh, off-peak 80 g; live/frontier work 80% on-peak; 100% execution of eligible scheduling. Demand reduction is fewer kWh at the natural hour mix. Scheduling is the additional effect on remaining HEA kWh. One EU passenger car ≈ 2 tCO2/year. On the inverted grid, scheduling changes sign.
6. Sensitivity and limitations
A low band (5 people × 10 operations × 0.6 Wh frontier × weaker replaceable share) and a high band (20 people × 20 operations × 2.5 Wh × draft-heavy displacement) bound the mid case by roughly an order of magnitude at each scale. The conversation-centric 108-client case reaches about 11 TWh net energy avoided and about 3.8 MtCO2 combined avoided per year — a few percent of a 500 TWh-class global datacenter total, or about 1.9 million EU-car years [1]. That is meaningful progress if the displacement mix is real. It is a ceiling for this model, not a sales plan.
6.1a Three demand stories at one million clients
Table 4 holds client count at 106 and compares three scenario bundles. Mid is today’s operator-day mix. Future keeps that staff day and adds business-plan visitor Q&A plus lean automation. Conversation-centric uses the support-site funnel and splits chat into frontier, legacy service, and induced shares. Volume, substitution shares, and schedulability all change; the energy coefficients, natural hour mix, scheduling execution, tariff, and carbon grid stay fixed.
| Mid | Future | Conversation-centric | |
|---|---|---|---|
| Ops / year (all clients) | 33 billion | 206 billion | 430 billion |
| Chat share of ops / of HEA kWh | 25% / 45% | 51% / 75% | 72% / 88% |
| Same-workload frontier → HEA electricity | 9.38 → 4.04 GWh | 121 → 30.2 GWh | 184 → 76.0 GWh |
| Net energy vs replaced frontier | 5.34 GWh | 91.1 GWh | 108 GWh |
| Energy save | 57% | 75% | 59% |
| Schedulable / executed / effective off-peak | 25% / 100% / 40% | 18% / 100% / 34% | 9% / 100% / 27% |
| Same-workload frontier → HEA footprint | 3.15 → 1.10 kt | 40.76 → 8.78 kt | 61.87 → 23.87 kt |
| Demand + scheduling = combined avoided | 1.80 + 0.26 = 2.05 kt | 30.60 + 1.39 = 31.99 kt | 36.33 + 1.67 = 38.00 kt |
| HEA footprint vs today’s mid frontier | 2.05 kt lower | 5.62 kt higher | 20.71 kt higher |
Table 4. One million clients. Natural live/frontier mix is 80% on-peak; scheduling execution is 100% of eligible automation. Future chat replaceable share is 65%. Conversation-centric chat is 35% frontier / 45% legacy / 20% induced. Combined avoided emissions compare each demand story with its own same-workload frontier counterfactual. The final row instead compares each HEA footprint with today’s smaller mid frontier footprint; future demand is therefore higher in absolute terms even while avoiding most of its counterfactual emissions.
6.1b Conversation-centric installed base
Table 5 holds the conversation-centric mix and scales only the installed base. Energy save stays 59%; combined carbon avoided includes both lower demand and eligible scheduling. At 10 million clients the yearly net energy is about 1.1 TWh and combined avoided emissions are about 380 ktCO2. At 100 million clients the line is about 10.8 TWh and 3.80 MtCO2, on the order of 1.9 million EU-car years. Legacy contacts remain a service count, not a second watt-hour column.
| Clients | HEA ops / year | HEA energy | Net energy | Flat net $ | Combined avoided CO2 | EU-car years |
|---|---|---|---|---|---|---|
| 100,000 | 43.0 billion | 7.60 GWh | 10.8 GWh | $1.08M | 3.80 kt | 1,900 |
| 1 million | 430 billion | 76.0 GWh | 108 GWh | $10.8M | 38.0 kt | 19,000 |
| 10 million | 4.30 trillion | 760 GWh | 1.08 TWh | $108M | 380 kt | 190,000 |
| 100 million | 43.0 trillion | 7.60 TWh | 10.8 TWh | $1.08B | 3.80 Mt | 1.9 million |
Table 5. Conversation-centric band only, one year. Natural live/frontier mix is 80% on-peak; all eligible automation is scheduled off-peak, producing a 27% effective HEA off-peak share. Energy save is 59% at every scale. One EU passenger car ≈ 2 tCO2/year. 10.8 TWh is about twelve 876 GWh mid-size datacenter-years, or about 2% of a 500 TWh datacenter class. Legacy contacts at these scales are not added to watt-hour or CO2 columns.
6.1c Infrastructure horizon: cumulative and Net Present CO2 Saving
Table 6 treats the conversation-centric architecture as a system that can keep serving if quality and the replaceable mix hold. Installed base is held constant — no adoption ramp and no assumption that frontier assistants themselves shrink. Cumulative is n × the yearly combined avoided line. NPCS discounts each future year at 7.6% [16]. At 100 million clients, the 10-year row is 38.0 MtCO2 cumulative and 25.96 Mt NPCS. A 25-year row is 95.0 Mt cumulative and 41.99 Mt NPCS. Those are same-workload counterfactual results, not absolute reductions below today or an adoption forecast.
| Horizon | 10M cumulative | 10M NPCS | 10M car-years | 100M cumulative | 100M NPCS | 100M car-years |
|---|---|---|---|---|---|---|
| 5 years | 1.90 Mt | 1.53 Mt | 0.95 million | 19.0 Mt | 15.33 Mt | 9.5 million |
| 10 years | 3.80 Mt | 2.60 Mt | 1.9 million | 38.0 Mt | 25.96 Mt | 19.0 million |
| 25 years | 9.50 Mt | 4.20 Mt | 4.75 million | 95.0 Mt | 41.99 Mt | 47.5 million |
Table 6. Conversation-centric band; constant installed base; r = 7.6%; end-of-year annuity. Cumulative energy at 100 million clients is 54 / 108 / 270 TWh for 5 / 10 / 25 years. Undiscounted electricity dollars at that scale are $5.4B / $10.8B / $27.0B (tariff, not a WACC NPV). A 17% discount rate remains available as a calculator sensitivity.
6.2 Conversation counts per client
Table 7 is the engagement view the energy columns cannot show. Legacy contacts are email, form, webpage, and phone contact-equivalents. They do not enter Tables 1–6 as frontier watt-hours.
| Mid | Future | Conversation-centric | |
|---|---|---|---|
| Visitor chat turns / client / month | 0 | 8,000 | 25,000 |
| All chat turns / client / month | 688 | 8,688 | 25,688 |
| Frontier-replaced turns / client / year | 2,475 | 67,763 | 107,888 |
| Legacy contacts / client / year | 0 | 0 | 138,713 |
| Legacy contacts at 1 million clients | 0 | 0 | 139 billion |
| Legacy contacts at 10 million clients | 0 | 0 | 1.39 trillion |
| Legacy contacts at 100 million clients | 0 | 0 | 13.9 trillion |
Table 7. Staff chat is the mid 10 × 15 × 25% slice (688 turns / month). Visitor volume is the HEA World business-plan Q&A surface: future uses an 8,000-turn/month order of magnitude; conversation-centric uses the Support funnel 10,000 uniques × 0.25 × 2.5 × 4. Each legacy-replaced turn is one contact-equivalent. Channel mix is not estimated. Contact totals scale with installed base; they are not a second CO2 column.
Limitations are first-class:
- Energy coefficients are scenario intensities, not metered GPU traces.
- Quality must hold. If users bounce from mini to a frontier assistant, the saving is fictional. Model qualification is part of the architecture, not an afterthought.
- Induced demand can erase net saving if HEA creates more operations than it displaces. The future case assumes most website Q&A would otherwise have been a frontier assistant. If visitors only chat because a widget exists, keep the mid replaceable share. The conversation-centric case already holds 20% of chat as induced.
- Legacy contacts are not an energy credit. Retiring email, forms, pages, and phone changes operations; it does not automatically change GPU watt-hours.
- Avoided future emissions are not automatically an absolute reduction from today. The paper reports the same-workload frontier counterfactual and the HEA footprint against today’s mid frontier footprint separately.
- Mailbox classification must not be multiplied by headcount. It belongs to the inbox.
- Jevons effects (cheaper inference, more use) are not modelled beyond the induced-demand share.
- Coding, creative, and open research frontier use are out of scope.
7. Discussion
The distinctive result is not “small models are cheaper.” It is that organization-specific work can be compiled so that the live model no longer has to be large. That is closer to classical compiler design than to prompt engineering. Token telemetry already used for billing and quota planning is the same telemetry a later measured paper would need: tokens per turn, model mix, and a diary of replaceable share on a real cohort.
For climate accounting, hourly carbon intensity on a named grid and year (Electricity Maps, Ember, or a transmission-system operator) should replace the three scenario pills before anyone treats Table 3 as an inventory line [14]. For commercial accounting, electricity dollars should not be compared with software invoices. For public-policy language, Goals 7, 9, 12, and 13 name the questions this paper is asking; they are not a score [15].
Conversation-centric engagement is the product thesis that is easy to smuggle into the energy thesis. A support-site HEA can replace a frontier assistant and a contact-centre loop. Only the first belongs in Tables 1–6. Table 7 keeps the second visible without adding it to ChatGPT watt-hours. At 100 million clients, combined avoided emissions are about 3.80 MtCO2/year against the same-workload frontier counterfactual, and 10-year NPCS is about 25.96 Mt [16]. The HEA footprint is still about 2.39 Mt/year, above today’s smaller 0.315 Mt mid frontier footprint. Both statements are true: the architecture avoids a large future counterfactual and demand growth still raises absolute emissions.
8. Conclusion
Compile once, serve many is an inference architecture. It downshifts live organizational work from frontier models to mini and nano calls, then moves eligible automation into cheaper and — on the right grid — cleaner hours. Those are separate effects. At 100 million clients, the modeled annual combined avoided emissions are about 0.205 Mt in mid, 3.20 Mt in future, and 3.80 Mt in conversation-centric; 10-year NPCS at 7.6% is about 1.40, 21.86, and 25.96 Mt. Future demand still raises the HEA footprint above today’s smaller workload, so the paper reports avoided counterfactual emissions and absolute change from today side by side. Legacy email, form, page, and phone contacts remain a service count, not an energy credit. The 2030 Agenda names the policy questions; it is not a product score. Client-count totals are sensitivity tables. The open calculator is the methods appendix that readers can argue with.
Companion calculator
The interactive model uses the same equations and defaults. Change people, operations per day, tariff, live hour mix, eligible scheduling execution, carbon grid, and NPCS horizon. It reports demand reduction, scheduling, combined avoided emissions, HEA footprint, and absolute change from today. No data leave the page.
References
- IEA (2025). Energy and AI. Paris: International Energy Agency. https://www.iea.org/reports/energy-and-ai
- Luccioni, A. S., Jernite, Y., & Strubell, E. (2024). Power Hungry Processing: Watts Driving the Cost of AI Deployment? ACM FAccT. https://doi.org/10.1145/3630106.3658542
- Patterson, D., Gonzalez, J., Le, Q., Liang, C., Munguia, L.-M., Rothchild, D., So, D., Texier, M., & Dean, J. (2021). Carbon emissions and large neural network training. arXiv:2104.10350.
- Strubell, E., Ganesh, A., & McCallum, A. (2019). Energy and policy considerations for deep learning in NLP. ACL.
- Schwartz, R., Dodge, J., Smith, N. A., & Etzioni, O. (2020). Green AI. Communications of the ACM, 63(12), 54–63.
- Luccioni, A. S., Viguier, S., & Ligozat, A.-L. (2023). Estimating the carbon footprint of BLOOM, a 176B parameter language model. Journal of Machine Learning Research, 24(253), 1–15.
- Dodge, J., Prewitt, T., Tiwari, R., et al. (2022). Measuring the carbon intensity of AI in cloud instances. ACM FAccT.
- Payen, N. (2026). Human-Enhanced Agents: The ownership layer of the agentic web. HEA World. https://hea-world.com/articles/human-enhanced-agents-white-paper.html
- de Vries, A. (2023). The growing energy footprint of artificial intelligence. Joule, 7(10), 2191–2194.
- Eurostat. Electricity prices for non-household consumers. https://ec.europa.eu/eurostat
- U.S. Energy Information Administration. Electric Power Monthly, industrial sector prices. https://www.eia.gov
- IPCC (2022). Climate Change 2022: Mitigation of Climate Change. Working Group III, Annex III, lifecycle greenhouse-gas intensities of electricity generation.
- Ember (2025). European Electricity Review. https://ember-energy.org
- Electricity Maps. Hourly carbon intensity methodology. https://www.electricitymaps.com
- United Nations General Assembly (2015). Transforming our world: the 2030 Agenda for Sustainable Development. A/RES/70/1. Official goal titles: Goal 7 Affordable and clean energy; Goal 8 Decent work and economic growth; Goal 9 Industry, innovation and infrastructure; Goal 12 Responsible consumption and production; Goal 13 Climate action. https://sdgs.un.org/2030agenda
- Payen, N. (2021). Why and how to use Net Present CO2 saving for smarter climate action decisions. https://npayen.com/articles/12.html (also LinkedIn). Discount rate = Paris-pathway annual cut; 2023 note records a 17% PwC update as a stricter sensitivity.
- UNEP (2019). Emissions Gap Report 2019. 7.6% annual reduction 2020–2030 for a 1.5 °C pathway. https://www.unep.org/resources/emissions-gap-report-2019
- Wu, C.-J., Raghavendra, R., Gupta, U., et al. (2022). Sustainable AI: Environmental implications, challenges and opportunities. MLSys.
- Hugging Face (2025). AI Energy Score. https://huggingface.co/blog/ai-energy-score
- OpenAI. API pricing (accessed 2026). Used only as a footnote on vendor invoices, not as energy. https://openai.com/api/pricing/
- Payen, N. (2026). Preparing HEA-World for AI Act Article 50. HEA World. https://hea-world.com/articles/preparing-for-ai-act-article-50.html
Appendix. Default mid-case parameters
10 people per client; 15 AI operations per person per day; 220 workdays; mix 25/40/20/15 (chat / classify / draft / other); replaceable shares 30/15/70/50; induced shares 40/20/15/25; frontier query 1.5 Wh; Foundry 40 Wh and USD 12 token-equivalent per HEA-year (electricity path uses Wh only); flat USD 0.10 / kWh; TOU USD 0.18 / 0.06; live/frontier on-peak share 80%; eligible scheduling execution 100%; carbon 400 / 80 gCO2/kWh.
Future case (calculator band future): same staff day; extra monthly ops per client 8,000 visitor chat turns, 4,000 classify, 400 queued drafts, 2,000 other lean jobs; chat replaceable 65%; architectural shiftability 0 / 1 / 0.35 / 0.70 (chat / classify / draft / other); Foundry 55 Wh. Visitor volume is an order-of-magnitude HEA World business-plan Q&A surface (104 monthly uniques at support-like engagement), not a metered cohort.
Conversation-centric case (calculator band centric): same staff day; extra monthly ops 25,000 visitor chat turns, 5,000 classify, 600 queued drafts, 2,500 other lean jobs; chat 35% frontier-replaceable / 45% legacy-replaceable / 20% induced; same architectural shiftability as future; Foundry 55 Wh. The 25,000 turns are the Support funnel 10,000 uniques × 0.25 × 2.5 × 4. Legacy contacts are not a second kWh column. Horizon / NPCS: 5, 10, 25 years; default r = 7.6% end-of-year (Payen 2021 / UNEP 2019); optional 17%. Source code of the browser model: /scripts/web/hea_energy_scenario.js (v1.4).
