Modern LLMs are data and hence energy hungry
Links to scaling
Read the following article in Nature on the impacts this is having in the Global South
🤔 That is why batching, caching, quantization and speculation are so important
🤔 Politeness costs tokens: should you be polite to your LLM?
Cost is mostly in GPU time
🎮 Practicals on estimating energy usage (and links to scaling) in LLMs
Queue management/SLO aware scheduling
Carbon aware computing: serving across geographical data centres
In-flight efficiency
📝 CEDAR paper GreenSys paper Amit More, Tarique Anwar, Poonam Yadav, 2026
Energy = Energy(prefill) + Energy(decode)