Modern LLMs are data and hence energy hungry
Links to scaling
Read the following article in Nature on the impacts this is having in the Global South
🎮 Practicals on estimating energy usage (and links to scaling) in LLMs
Queue management/SLO aware scheduling
KV cache management
The Sunk Carbon Fallacy Bashir et al.
Carbon aware computing: serving across geographical data centres
In-flight efficiency
📝 CEDAR paper GreenSys paper Amit More, Tarique Anwar, Poonam Yadav, 2026
Energy = Energy(prefill) + Energy(decode)