How Do Prompt Variations Affect Energy Consumption in On-Device LLMs?

arXiv:2609.01798 · cs.CL · Submitted 2026-09-01 · Read on arXiv

cs.CL

Submitted: 2026-09-01

Updated: 2026-09-01

Comments: Accepted to the EMNLP 2026 Main Conference; camera-ready version

Code: https://github.com/confident-ai/deepeval

Project page: https://amai-gsu.github.io/PromptProperty

License: http://creativecommons.org/licenses/by/4.0/

The gist: Large language models (LLMs) are increasingly deployed on mobile devices, making energy efficiency a key deployment constraint, yet the energy impact of prompt design remains underexplored.

Terminology

Abstract

Large language models (LLMs) are increasingly deployed on mobile devices, making energy efficiency a key deployment constraint, yet the energy impact of prompt design remains underexplored. This paper aims to understand how two prompt properties, cognitive load and phrasing pattern, shape the energy behavior of on-device LLM inference. We conduct a broad empirical study covering prompt properties, datasets, models, and devices, with phase-level profiling that separates prefill and decode energy. We find that cognitive load primarily affects the energy cost per token, while phrasing pattern affects energy largely through token usage. Our energy-quality analysis further shows that prompt design reshapes the attainable frontier differently across models, highlighting the need for model-aware prompt design in energy-efficient on-device LLM inference. Code, datasets, and scripts are available at https://amai-gsu.github.io/PromptProperty/.

Sources

Related papers