Surprisal Theory is Tautological (without Rational Grounding)
summary
The gist
Surprisal theory is critiqued for being tautological without rational grounding, suggesting that its falsifiability requires restricting the language model to one grounded in non-empirically
In short
Surprisal theory claims that processing difficulty is determined by a language model's surprisal, but this is unfalsifiable because any difficulty pattern can be explained by constructing a specific language model to fit it. The paper argues this requires restricting the set of possible models (Q) based on cognitive principles rather than just empirical data, suggesting a need for rational grounding to make the theory testable.
Key concepts
- Tautology Argument
- The core critique is that surprisal theory is circular; for any observed processing difficulty, one can always build a language model whose surprisal is an affine function of it. This means the relationship between difficulty and surprisal is not unique to a single model but can be satisfied by countless models, rendering the original claim unfalsifiable without external constraints.
- Cognitivist Assumption
- To solve the tautology, one must restrict language models (Q) based on internal cognitive principles, such as memory limits or processing goals. This assumption suggests that the relevant model should approximate a theoretical construct derived from how a human mind actually processes information, moving the grounding from reading-time data to observable cognitive constraints.
- Scaling Implication
- This idea predicts that as language models get better at approximating the training corpus (pC), their surprisal values should improve and better correlate with actual processing difficulty. However, empirical evidence shows this scaling breaks down, suggesting that simply fitting a larger model to the data does not necessarily lead to a more accurate cognitive representation.
Terminology used across episodes
This episode discusses
- Surprisal Theory is Tautological (without Rational Grounding) · Paper Radio
- surprisal is Not a Theory · Paper Radio
The paper
Surprisal Theory is Tautological (without Rational Grounding) · Read on arXiv
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "Surprisal Theory is Tautological (without Rational Grounding)".
Tom: Surprisal theory is critiqued for being tautological without rational grounding, suggesting that its falsifiability requires restricting the language model to one grounded in non-empirically motivated cognitive principles.
Jane: First, who's behind it and why it matters.
Title and authors: Tom: So, Ryan Cotterell’s paper, "Surprisal Theory is Tautological (without Rational Grounding)," basically argues that the core idea of surprisal theory—that difficulty relates to how surprising a linguistic unit is under some model—is unfalsifiable without adding serious constraints.
Jane: That means that for any observed pattern of how hard it is to process something, you can just construct a language model whose surprisal matches it perfectly, which the paper calls a tautology <ref:2607.21574#pg1>.
Lu: That's the big conceptual hurdle they are pointing out; it shows that without external grounding, the theory doesn't actually tell us much about human processing itself.
Meng: So, the authors are suggesting that to make this theory useful again, we have to narrow down the set of possible language models into a specific family called Q.
Lalam: Exactly! And they highlight that picking which model from Q is the second major part of their grounding problem, which is where we need real cognitive input instead of just fitting data.
The paper's summary: Tom: To summarize what the authors laid out in this paper, they argue that surprisal theory claims processing difficulty is an affine function of contextual surprisal under some language model q, but they show that any non-negative difficulty measure d can be matched by a language model whose surprisal is an affine function of it <ref:2607.21574#pg0>.
Jane: They point out that this is because the theory leaves the language model q as a free variable, meaning any observed pattern of difficulty can just be explained by positing a specific q whose surprisal matches it exactly.
Lu: The paper really zeroes in on the fact that without an independent characterization of pH that isn't based on behavior, we have no way to tell which model is actually relevant to human cognition.
Meng: So, the authors are essentially saying that if we don't define what q should be based on cognitive principles instead of just fitting corpus data, the theory loses its predictive power.
Lalam: That’s a key point for us; it moves the focus from simply measuring correlation to understanding the actual structure of human language processing.
The paper's improvements: Tom: So where does this paper suggest we go from here? The authors propose that breaking this tautology requires a rationalist intervention, meaning we need to choose a specific family Q and select a model q based on cognitive principles rather than just behavioral data <ref:2607.21574#pg2>.
Jane: They suggest two main paths for this grounding: either using lossy-context surprisal where we replace the full context with something motivated by memory decay, or using corpora that are grounded in language acquisition research.
Lu: The idea of lossy context is really cool because it suggests modeling human memory constraints directly into the model's architecture rather than just letting it learn from massive datasets.
Meng: If we look at implementation, this means designing AI systems that explicitly account for resource allocation policies—like memory limits—instead of relying solely on what the training data implies about processing difficulty.
Lalam: I think focusing on those structural commitments, like restricting Q to PCFGs if we care about syntax, gives us a much more stable and testable framework than letting it be any transformer.
Conclusion: Tom: So to wrap up this discussion on "Surprisal Theory is Tautological (without Rational Grounding)," the main point is that without constraining the language model q, surprisal theory remains unfalsifiable because we can always construct a matching model <ref:2607.21574#pg0>.
Jane: The authors are urging us toward a rationalist intervention where we select models based on cognitive aspects of the human brain rather than just observing reading-time data.
Lu: This opens up huge possibilities for creating more theoretically sound language models that actually mirror how language is processed incrementally, which is what I find really exciting about this work.
Meng: For practical application, it means we need to build mechanisms into our AI that respect known cognitive limits, like memory decay or working memory capacity.
Lalam: Ultimately, the paper shows us that the next step isn't just fitting more data; it's about building a language model whose internal structure reflects a more accurate understanding of how human cognition actually operates under resource limitations.
More episodes
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language
- 2508.08833-An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems
- 2405.04118-Policy Learning with a Language Bottleneck