Intelligence per Watt: Measuring Intelligence Efficiency of Local AI
cs.DC, cs.AI, cs.CL, cs.LG
Submitted: 2025-11-11
Updated: 2026-10-01
Code: https://github.com/HazyResearch/intelligence-per-watt
Project page: https://hazyresearch.stanford.edu/intelligence-per-watt
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large language model (LLM) queries are predominantly processed by frontier models in centralized cloud infrastructure.
Terminology
Abstract
Large language model (LLM) queries are predominantly processed by frontier models in centralized cloud infrastructure. Demand growth strains this paradigm faster than providers can scale. Two advances create an opportunity to rethink it: small, local LMs (<=20B active parameters) now achieve competitive performance to frontier models on many tasks, and local accelerators (e.g., Apple M4 Max) can host these models at interactive latencies. This raises the question: can local inference viably redistribute demand from centralized infrastructure? This requires measuring both whether local LMs can accurately answer real-world queries and whether they can do so efficiently on power-constrained devices (e.g., laptops). We propose intelligence per watt (IPW), task accuracy per unit of power, as a unified metric for the capability and efficiency of local inference across model-accelerator configurations. We evaluate 20+ state-of-the-art local LMs, 8 hardware accelerators (local and cloud), and 1M real-world single-turn chat and reasoning queries. For each query, we measure accuracy (local LM win rate against frontier models), energy, latency, and power. We find three key results. First, local LMs successfully answer 88.7% of these queries, with accuracy varying by domain. Second, longitudinal analysis from 2023-2025 shows IPW improved 5.3x, driven by both algorithmic and accelerator advances, with locally-serviceable query coverage rising from 23.2% to 71.3%. Third, local accelerators achieve at least 1.4x lower IPW than cloud accelerators running identical models, revealing significant headroom for local accelerator optimization. These findings demonstrate that local inference can meaningfully redistribute demand from centralized infrastructure for a substantial subset of queries, with IPW serving as the critical metric for tracking this transition.
Sources
- gpt-oss-120b & gpt-oss-20b Model Card
- Carbontracker: Tracking and Predicting the Carbon Footprint of Training Deep Learning Models
- Constitutional AI: Harmlessness from AI Feedback
- FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- DeepSeek-V3 Technical Report
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- WildVis: Open Source Visualizer for Million-Scale Chat Logs in the Wild
- Early ChatGPT User Portrait through the Lens of Data
- Energy Considerations of Large Language Model Inference and Efficiency Optimizations
- Gemma 3 Technical Report
- The Llama 3 Herd of Models
- Which Economic Tasks are Performed with AI? Evidence from Millions of Claude Conversations
- Distilling the Knowledge in a Neural Network
- RouterBench: A Benchmark for Multi-LLM Routing System
- RouterEval: A Comprehensive Benchmark for Routing LLMs to Explore Model-level Scaling Up in LLMs
- Mixtral of Experts
- SLED: A Speculative LLM Decoding Framework for Efficient Edge Serving
- SpecInfer: Accelerating Generative Large Language Model Serving with Tree-based Speculative Inference and Verification
- Minions: Cost-efficient Collaboration Between On-device and Cloud Language Models
Related papers
- iScheduler: Reinforcement Learning-Driven Continual Optimization for Large-Scale Resource Investment Problems
- SAMM: Sharded Automated Market Maker
- InferScale: GPU-Native KV Injection for Personalized LLM Serving
- Vigil: Accountable Liveness against Selective Silence
- Steelhead: Interleaving Partially Synchronous and Asynchronous Commit Rules on a Shared DAG
- Pushing CPU Speech Synthesis to the Wall: Extreme Inference Tuning under Serverless Architecture and Billing