surprisal is Not a Theory
cs.CL
Submitted: 2026-07-22
Updated: 2026-09-15
Project page: https://osf.io/n8kb3/overview
License: http://creativecommons.org/licenses/by-sa/4.0/
The gist: Surprisal Theory is often characterized as a computational-level explanation per (Marr, 1982).
Terminology
Abstract
Surprisal Theory is often characterized as a computational-level explanation per (Marr, 1982). We argue in this work that, even though a computational level narrative has been used to support "representation-agnostic research" within computational psycholinguistics, the movement toward black box systems embodied by large language models (LLMs) does not exempt modelers using the surprisal metric from the representational decisions required by computational-level characterizations. In fact, we argue that the uncritical use of LLM-surprisal obfuscates the representational and algorithmic-level commitments of different models. In three analyses, we show that the choice of algorithm and model architecture play significant roles in the computation of language model probabilities. We advise that researchers who wish to test Surprisal Theory re-evaluate the practice of treating large language model probabilities as interchangeable
Sources
- Trajectory Dynamics in Language Model Hidden States Predict Human Processing Costs Beyond Surprisal
- On the Opportunities and Risks of Foundation Models
- RoBERTa: A Robustly Optimized BERT Pretraining Approach
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering