FrugalPrompt: Reducing Contextual Overhead in Large Language Models via Token Attribution
cs.CL
Submitted: 2025-10-18
Updated: 2026-09-19
Project page: https://frugalprompt.github.io
License: http://creativecommons.org/licenses/by/4.0/
The gist: Human communication heavily relies on laconism and inferential pragmatics, allowing listeners to successfully reconstruct rich meaning from sparse, telegraphic speech.
Terminology
Abstract
Human communication heavily relies on laconism and inferential pragmatics, allowing listeners to successfully reconstruct rich meaning from sparse, telegraphic speech. In contrast, large language models (LLMs) owe much of their stellar performance to expansive input contexts, yet such verbosity inflates monetary costs, carbon footprint, and inference-time latency. This overhead manifests from the redundant low-utility tokens present in typical prompts, as only a fraction of tokens typically carries the majority of the semantic weight. Inspired by the aforementioned cognitive psycholinguistic processes, we address this inefficiency by introducing FrugalPrompt, a novel prompt compression framework for LLMs, which retains only the most semantically significant tokens. Leveraging two state-of-the-art token attribution methods, GlobEnc and DecompX, we assign salience scores to every token in an input sequence, rank them to retain the top-k% tokens, and obtain a sparse frugalized prompt. We establish the theoretical stability of our approach and provide strong empirical results across a suite of four NLP tasks to study the trade-off between the portion of retained tokens and performance. Experimental findings across retention settings reveal asymmetric performance patterns that suggest potential task contamination effects. We posit that our work contributes to a more nuanced understanding of LLM behavior in performance-efficiency trade-offs and delineates the boundary between tasks tolerant of contextual sparsity and those requiring exhaustive context.
Sources
- What Does BERT Look At? An Analysis of BERT's Attention
- Training Verifiers to Solve Math Word Problems
- SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression
- Mamba: Linear-Time Sequence Modeling with Selective State Spaces
- Efficient Prompting Methods for Large Language Models: A Survey
- Do Attention Heads in BERT Track Syntactic Dependencies?
- Generating Long Sequences with Sparse Transformers
- Attention is not Explanation
- Agent Laboratory: Using LLM Agents as Research Assistants
- Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training
- Towards Greener LLMs: Bringing Energy-Efficiency to the Forefront of LLM Inference
- Efficient Large Language Models: A Survey
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering