Summary
- Report energy by model task, hardware, location, time and useful output, including idle and cooling overhead.
- Efficiency gains should be checked against growth in usage, grid constraints and water or backup-power impacts.
AI electricity demand comes from dense computation, memory movement, networking and cooling, with very different profiles between training and serving. Better chips can lower energy per request while total demand still rises through scale. The next useful evidence is an auditable meter trail from workload scheduler to facility and grid interval, paired with service quality. Decisions improve when operators can shift, resize or stop a workload based on its real marginal cost rather than an annual average.


