How Do Prompt Variations Affect Energy Consumption in On-Device LLMs?
cs.CL
Submitted: 2026-09-01
Updated: 2026-09-01
Comments: Accepted to the EMNLP 2026 Main Conference; camera-ready version
Code: https://github.com/confident-ai/deepeval
Project page: https://amai-gsu.github.io/PromptProperty
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large language models (LLMs) are increasingly deployed on mobile devices, making energy efficiency a key deployment constraint, yet the energy impact of prompt design remains underexplored.
Terminology
Abstract
Large language models (LLMs) are increasingly deployed on mobile devices, making energy efficiency a key deployment constraint, yet the energy impact of prompt design remains underexplored. This paper aims to understand how two prompt properties, cognitive load and phrasing pattern, shape the energy behavior of on-device LLM inference. We conduct a broad empirical study covering prompt properties, datasets, models, and devices, with phase-level profiling that separates prefill and decode energy. We find that cognitive load primarily affects the energy cost per token, while phrasing pattern affects energy largely through token usage. Our energy-quality analysis further shows that prompt design reshapes the attainable frontier differently across models, highlighting the need for model-aware prompt design in energy-efficient on-device LLM inference. Code, datasets, and scripts are available at https://amai-gsu.github.io/PromptProperty/.
Sources
- SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model
- From Prompts to Power: Measuring the Energy Footprint of LLM Inference
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- The Llama 3 Herd of Models
- Distilling the Knowledge in a Neural Network
- The Price of Prompting: Profiling Energy Use in Large Language Models Inference
- MobileAIBench: Benchmarking LLMs and LMMs for On-Device Use Cases
- Gemma 2: Improving Open Language Models at a Practical Size
- A Survey of Prompt Engineering Methods in Large Language Models for Different NLP Tasks
- PowerInfer-2: Fast Large Language Model Inference on a Smartphone
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering