Superficial Beliefs in LLM Decision-Making
summary
The gist
This paper investigates whether large language models (LLMs) possess a systematic underlying decision structure or merely "imitate the language of belief." By comparing an LLM's actual choices
In short
The paper "Superficial Beliefs in LLM Decision-Making" by Gabriel Freedman and Francesca Toni examines how Large Language Models' verbal explanations often mismatch their actual decision-making logic. While behavioral models can predict choices with 80.4% accuracy, the attributes models claim are important only align with their actions 61% of the time.
Key concepts
- Superficial Beliefs
- This refers to a disconnect where Large Language Models make decisions based on an underlying logic but provide verbal explanations that do not match what actually drove the choice. The models may provide plausible-sounding reasons that are inconsistent with their actual internal decision-making patterns.
- Score-based Judge
- Instead of asking a model to name one single most important attribute for a decision, this method asks the model to provide a decisiveness score for every factor involved. This approach is more stable and reproducible than relying on direct verbal explanations.
Terminology used across episodes
This episode discusses
- Superficial Beliefs in LLM Decision-Making · Paper Radio
- Reasoning Models Don't Always Say What They Think
- Alignment Revisited: Are Large Language Models Consistent in Stated and Revealed Preferences?
- Language Models (Mostly) Know What They Know
- Ministral 3
- OpenAI GPT-5 System Card
- Qwen3 Technical Report
The paper
Superficial Beliefs in LLM Decision-Making · Read on arXiv
Imperial College London
We ask whether large language models (LLMs) merely imitate rationales when choosing between two options, or whether their choices reflect a systematic underlying decision structure. Using synthetic binary decision settings in which models choose between profiles defined by graded attributes, we compare the attribute a model says mattered most with the attribute that best explains its choice under a behavioural model fit to prior decisions. The behavioural model predicts held-out choices well, showing that model behaviour is systematically related to the visible attributes rather than being random. However, direct self-reports and a separate score-based judge recover the behaviourally inferred driver only partially. The resulting picture is neither one of arbitrary behaviour nor one of fully articulated belief - outputs are structured enough to support prediction, but explicit reasons track the recovered driver only imperfectly. This qualitative pattern persists across prompt-order and sampling perturbations, alternative behavioural models, targeted occlusion analyses, and structurally varied decision settings. We interpret this as evidence for ``superficial belief'' in LLM decision-making: models behave as if guided by probabilistic local priorities over attributes, while having only limited verbal access to the attributes that drive their decisions.
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Superficial Beliefs in LLM Decision-Making".
Jane: The paper was written by Gabriel Freedman and Francesca Toni from Imperial College London.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Title: Tom: Welcome to the show! Today we are looking at a fascinating new paper titled "Superficial Beliefs in LLM Decision-Making" by Gabriel Freedman and Francesca Toni from Imperial College London.
Jane: It’s a title that really makes you stop and think, Tom. It suggests that these models might be performing a kind of intellectual masquerade.
Tom: That is exactly the vibe I got from it! Jane, how would you explain what they mean by "superficial belief" to someone who isn't a researcher?
Jane: Think of it like this: imagine someone makes a decision based on their gut feeling, but then they invent a logical-sounding reason afterward just to satisfy you. The paper argues that LLMs might be doing the same thing, where their actual choices follow a consistent logic, but the verbal reasons they give don't actually match what drove the decision.
Lu: That sounds like we might be looking at a kind of digital subconscious where the real decisions happen! It is as if there is a whole layer of logic that exists but simply lacks the vocabulary to describe itself to us. We could eventually see models that possess incredibly deep reasoning without ever needing to explain it in human language.
Tom: A digital subconscious? That’s a wild way to put it, Lu!
Meng: It sounds like a massive headache for anyone trying to build reliable systems in the real world. If I am deploying an AI into a hospital or a bank, I cannot just take its word for it. I need to know that the "why" behind a decision is actually the truth, not just some plausible-sounding story the model generated to be polite.
Jane: You're right, Meng, and that's why this research is so vital for safety and trust.
Lalam: This could fundamentally change how our culture perceives intelligence. We have always tied "thinking" to the ability to explain oneself, but if models can act rationally without being able to articulate it, we might have to broaden our definition of a thinking entity.
Tom: That is a huge concept to wrap our heads around.
Jane: It really is, and we need to look at the actual data they collected to see how big this gap really is.
Summary: Tom: We have established the concept, but let's look at how Freedman and Toni actually proved this phenomenon exists.
Jane: They used these very structured synthetic tasks where models had to choose between two options based on four different attributes, like efficacy or safety, across themes like drugs and software.
Tom: And the results showed a really striking disconnect between what the models did and what they said, didn't they?
Jane: They certainly did. The researchers found that a behavioral model could predict the models' actual choices eighty point four percent of the time, but when they looked at which attribute the model *said* was important, that agreement dropped to only sixty-one point zero percent.
Meng: That gap is exactly what worries me from an engineering standpoint. You can have a system that is right eighty percent of the time, but if it's giving you the wrong reasons for those successes, you're essentially flying blind. If a model picks a drug because of safety but claims it picked it for efficacy, I can't use that to audit the system for errors.
Lu: It is almost like the model is following a hidden map that it can see, but it's trying to describe the journey using a completely different set of directions! The underlying logic is there and it's systematic, but the language layer is just disconnected from those internal weights.
Tom: So the logic is real, even if the explanation isn't?
Lu: Precisely! The model isn't just being random; it’s following a pattern that we can actually predict, even if the model itself can't verbalize that pattern accurately.
Lalam: It reminds me of how humans operate when we have an intuition or a "gut feeling." We often act on patterns our brains have recognized before our conscious mind has even had a chance to form a sentence.
Jane: That is such a helpful way to frame it, Lalam.
Tom: If we can't trust their words, how do we actually peek under the hood?
Improvements: Jane: That is the big question, and the paper actually points toward some better ways to probe these models.
Tom: They suggest using something called a "score-based judge" instead of just asking for a direct response, right?
Jane: Exactly. Instead of asking the model to name one single most important attribute, we ask it to give a decisiveness score for every factor involved.
Tom: So rather than getting one potentially fake reason, we get a breakdown of how much weight it assigns to everything?
Jane: Yes, and while that doesn't fix everything perfectly, the research shows it is much more stable and reproducible than just asking for a verbal explanation.
Lu: I wonder if we could take that even further with some truly creative interfaces! Imagine a dashboard where you don't read text at all, but instead see real-time probability maps or shifting weight visualizations as the AI processes a choice. We could turn that "superficial" layer into something visual and transparent.
Meng: That would be much more useful for my team than parsing unreliable text logs. If we can monitor those actual attribute weights, we can catch when a model is starting to drift from its intended logic during a live run. It turns a black box into something we can actually observe and control in real-time.
Jane: It really would move us toward much more objective debugging.
Lalam: I think this also shifts the goal of human-AI interaction. Instead of trying to force machines to speak like humans, we might learn to interpret their behavior through these structured patterns. We can build a culture where we value the consistency of an entity's actions as much as its ability to tell a good story.
Tom: That is a very profound way to look at the future of how we live alongside these systems.
Jane: It really is, and it's clear that there's still so much more to learn about these hidden drivers.
Conclusion: Tom: It has been an incredible session looking at "Superficial Beliefs in LLM Decision-Making."
Jane: We have certainly learned that what these models do and what they say are not always on the same page.
Lu: I am still buzzing about the idea of those visual weight interfaces! Being able to see the "why" through shifting patterns instead of just words could change everything about how we design software.
Meng: And I'll be thinking about how to build better testing frameworks that don't rely on these unfaithful verbal explanations. We need to focus on the behavioral data if we want to build truly safe and reliable AI.
Lalam: I see a future where we move past the illusion of conversation and learn to respect the mathematical patterns that actually drive these systems. It will lead to a much more authentic way of coexisting with technology.
Tom: Thanks for joining us, everyone!
Jane: We'll be back right after the break to talk about the massive energy costs of training these enormous models!
More episodes
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language