Is my model perplexed for the right reason? Contrasting LLMs' Benchmark Behavior with Token-Level Perplexity
cs.CL
Submitted: 2026-03-31
Updated: 2026-09-13
Code: https://github.com/sandropezzelle/Semantic_match_DUST
License: http://creativecommons.org/licenses/by/4.0/
The gist: Standard evaluations of Large language models (LLMs) focus on task performance, offering limited insight into whether correct behavior reflects appropriate underlying mechanisms and risking
Terminology
Abstract
Standard evaluations of Large language models (LLMs) focus on task performance, offering limited insight into whether correct behavior reflects appropriate underlying mechanisms and risking confirmation bias. We introduce a simple, principled interpretability framework based on token-level perplexity to test whether models rely on linguistically relevant cues. By comparing perplexity distributions over minimal sentence pairs differing in one or a few `pivotal' tokens, our method enables precise, hypothesis-driven analysis without relying on unstable feature-attribution techniques. Experiments on controlled linguistic benchmarks with several open-weight LLMs show that, while linguistically important tokens influence model behavior, they never fully explain perplexity shifts, revealing that models rely on heuristics other than the expected linguistic ones.
Sources
- TokenButler: Token Importance is Predictable
- Fixing confirmation bias in feature attribution methods via semantic match
- Perplexed: Understanding When Large Language Models are Confused
- Token-Level Adversarial Prompt Detection Based on Perplexity Measures and Contextual Information
- Mistral 7B
- Not all tokens are created equal: Perplexity Attention Weighted Networks for AI generated text detection
- Qwen2.5 Technical Report
- Gemma 3 Technical Report
- LLMs with Industrial Lens: Deciphering the Challenges and Prospects -- A Survey
- BLiMP: The Benchmark of Linguistic Minimal Pairs for English
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering