Intelligence per Watt: Measuring Intelligence Efficiency of Local AI
Summary
The paper introduces intelligence per watt (IPW) as a metric to evaluate the efficiency of local LMs on power-constrained devices and reports results across 20+ local models and 8 accelerators. It shows IPW improvements over time and demonstrates that local inference can redistribute demand from centralized cloud infrastructure for a substantial subset of queries.