Superficial Beliefs in LLM Decision-Making

arXiv:2606.11016 · cs.AI · Submitted 2026-06-09 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Superficial Beliefs in LLM Decision-Making".

Jane: The paper was written by Gabriel Freedman and Francesca Toni from Imperial College London.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title: Tom: Welcome to the show! Today we are looking at a fascinating new paper titled "Superficial Beliefs in LLM Decision-Making" by Gabriel Freedman and Francesca Toni from Imperial College London.

Jane: It’s a title that really makes you stop and think, Tom. It suggests that these models might be performing a kind of intellectual masquerade.

Tom: That is exactly the vibe I got from it! Jane, how would you explain what they mean by "superficial belief" to someone who isn't a researcher?

Jane: Think of it like this: imagine someone makes a decision based on their gut feeling, but then they invent a logical-sounding reason afterward just to satisfy you. The paper argues that LLMs might be doing the same thing, where their actual choices follow a consistent logic, but the verbal reasons they give don't actually match what drove the decision.

Lu: That sounds like we might be looking at a kind of digital subconscious where the real decisions happen! It is as if there is a whole layer of logic that exists but simply lacks the vocabulary to describe itself to us. We could eventually see models that possess incredibly deep reasoning without ever needing to explain it in human language.

Tom: A digital subconscious? That’s a wild way to put it, Lu!

Meng: It sounds like a massive headache for anyone trying to build reliable systems in the real world. If I am deploying an AI into a hospital or a bank, I cannot just take its word for it. I need to know that the "why" behind a decision is actually the truth, not just some plausible-sounding story the model generated to be polite.

Jane: You're right, Meng, and that's why this research is so vital for safety and trust.

Lalam: This could fundamentally change how our culture perceives intelligence. We have always tied "thinking" to the ability to explain oneself, but if models can act rationally without being able to articulate it, we might have to broaden our definition of a thinking entity.

Tom: That is a huge concept to wrap our heads around.

Jane: It really is, and we need to look at the actual data they collected to see how big this gap really is.

Summary: Tom: We have established the concept, but let's look at how Freedman and Toni actually proved this phenomenon exists.

Jane: They used these very structured synthetic tasks where models had to choose between two options based on four different attributes, like efficacy or safety, across themes like drugs and software.

Tom: And the results showed a really striking disconnect between what the models did and what they said, didn't they?

Jane: They certainly did. The researchers found that a behavioral model could predict the models' actual choices eighty point four percent of the time, but when they looked at which attribute the model *said* was important, that agreement dropped to only sixty-one point zero percent.

Meng: That gap is exactly what worries me from an engineering standpoint. You can have a system that is right eighty percent of the time, but if it's giving you the wrong reasons for those successes, you're essentially flying blind. If a model picks a drug because of safety but claims it picked it for efficacy, I can't use that to audit the system for errors.

Lu: It is almost like the model is following a hidden map that it can see, but it's trying to describe the journey using a completely different set of directions! The underlying logic is there and it's systematic, but the language layer is just disconnected from those internal weights.

Tom: So the logic is real, even if the explanation isn't?

Lu: Precisely! The model isn't just being random; it’s following a pattern that we can actually predict, even if the model itself can't verbalize that pattern accurately.

Lalam: It reminds me of how humans operate when we have an intuition or a "gut feeling." We often act on patterns our brains have recognized before our conscious mind has even had a chance to form a sentence.

Jane: That is such a helpful way to frame it, Lalam.

Tom: If we can't trust their words, how do we actually peek under the hood?

Improvements: Jane: That is the big question, and the paper actually points toward some better ways to probe these models.

Tom: They suggest using something called a "score-based judge" instead of just asking for a direct response, right?

Jane: Exactly. Instead of asking the model to name one single most important attribute, we ask it to give a decisiveness score for every factor involved.

Tom: So rather than getting one potentially fake reason, we get a breakdown of how much weight it assigns to everything?

Jane: Yes, and while that doesn't fix everything perfectly, the research shows it is much more stable and reproducible than just asking for a verbal explanation.

Lu: I wonder if we could take that even further with some truly creative interfaces! Imagine a dashboard where you don't read text at all, but instead see real-time probability maps or shifting weight visualizations as the AI processes a choice. We could turn that "superficial" layer into something visual and transparent.

Meng: That would be much more useful for my team than parsing unreliable text logs. If we can monitor those actual attribute weights, we can catch when a model is starting to drift from its intended logic during a live run. It turns a black box into something we can actually observe and control in real-time.

Jane: It really would move us toward much more objective debugging.

Lalam: I think this also shifts the goal of human-AI interaction. Instead of trying to force machines to speak like humans, we might learn to interpret their behavior through these structured patterns. We can build a culture where we value the consistency of an entity's actions as much as its ability to tell a good story.

Tom: That is a very profound way to look at the future of how we live alongside these systems.

Jane: It really is, and it's clear that there's still so much more to learn about these hidden drivers.

Conclusion: Tom: It has been an incredible session looking at "Superficial Beliefs in LLM Decision-Making."

Jane: We have certainly learned that what these models do and what they say are not always on the same page.

Lu: I am still buzzing about the idea of those visual weight interfaces! Being able to see the "why" through shifting patterns instead of just words could change everything about how we design software.

Meng: And I'll be thinking about how to build better testing frameworks that don't rely on these unfaithful verbal explanations. We need to focus on the behavioral data if we want to build truly safe and reliable AI.

Lalam: I see a future where we move past the illusion of conversation and learn to respect the mathematical patterns that actually drive these systems. It will lead to a much more authentic way of coexisting with technology.

Tom: Thanks for joining us, everyone!

Jane: We'll be back right after the break to talk about the massive energy costs of training these enormous models!

Imperial College London

cs.AI

Submitted: 2026-06-09

Updated: 2026-09-15

Comments: Published as a conference paper at COLM 2026

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 76/100

The gist: This paper investigates whether large language models (LLMs) possess a systematic underlying decision structure or merely "imitate the language of belief." By comparing an LLM's actual choices

Key concepts

Superficial Beliefs
This refers to a disconnect where Large Language Models make decisions based on an underlying logic but provide verbal explanations that do not match what actually drove the choice. The models may provide plausible-sounding reasons that are inconsistent with their actual internal decision-making patterns.
Score-based Judge
Instead of asking a model to name one single most important attribute for a decision, this method asks the model to provide a decisiveness score for every factor involved. This approach is more stable and reproducible than relying on direct verbal explanations.

Terminology

Summary

This paper investigates whether large language models (LLMs) possess a systematic underlying decision structure or merely imitate the language of belief. By comparing an LLM's actual choices against its explicit justifications, the researchers aim to determine if model reasoning is loosely connected to what actually drove the choice. This work is critical for understanding whether LLM outputs can be treated as guided by belief-like priorities or if they are simply arbitrary linguistic imitations.

The Benchmark Design

The researchers developed a synthetic benchmark of binary decision problems where models choose between two profiles defined by four graded attributes (low, medium, or high). To ensure the tasks involved genuine trade-offs, they only included pairs where each profile was better on at least one attribute. The study utilizes three substantive themes—drugs, policy, and software—each with a specific objective like optimising 5-year overall patient outcome. To validate that the models were not selecting attributes randomly, the researchers included two control themes where attributes were chosen to be irrelevant to the objective, such as Packaging Symmetry or Label Border Thickness. The benchmark is constructed using:

  • Four fixed global orders for attribute display.

  • Multiple prompt variants and repeated samples per prompt.

  • A balanced distribution of profiles appearing as both Option A and Option B.

Methodological Framework

The study employs an operational methodology to compare three different ways of identifying decision drivers:

  1. Direct Report: The model is prompted to choose exactly one option and then name the single most important attribute for that choice.

  2. Score-based Judge: The model provides a decisiveness score in [0, 1] for each of the four attributes, which are then aggregated using a signed-sum function to predict a choice.

  3. Behavioural Model: A binomial logistic regression is fitted to the models' choices to reconstruct local decision priorities over attributes. This identifies a revealed driver—the attribute that most strongly supports the side the model actually chose.

Key Findings and Results

The results support a middle position regarding LLM cognition. The behavioural model predicts held-out choices with high accuracy, matching the model's own choice 80.4% of the time on aggregate. However, at the attribute level, agreement is much weaker; direct reports and score-based judges recover the behaviourally inferred driver only partially, with rates of 61.0% and 61.3%, respectively. This suggests that while model behavior is systematically related to the visible attributes, their explicit reasons track the recovered driver only imperfectly. The authors interpret this as evidence for superficial belief: models behave as if guided by probabilistic local priorities over attributes while maintaining only limited verbal access to the actual drivers.

Robustness and Validation

The qualitative pattern of superficial belief is shown to be robust across several dimensions. The researchers found that:

  • Reproducibility and recovery come apart, meaning that while score-based outputs can be more reproducible than direct responses, they do not necessarily provide closer correspondence to the underlying decision structure.

  • Targeted occlusion analyses (using drop or equalise interventions) confirm that attributes ranked higher by the behavioural model produce larger disruptions in both choices and stated attributes.

  • The dissociation between choice and stated reason persists across different model families, prompt-order perturbations, and even a structurally different six-attribute decision setting.

Improvements for AI systems

1. Implementation of Parallel Behavioral Auditing Layers (Shadow Modeling)

  • Improvement: Integrate a non-linguistic, mathematical shadow model (e.g., a binomial logistic regression layer) that runs in parallel with the LLM during inference. This layer continuously fits the model's discrete choices to an attribute-difference framework to reconstruct its revealed decision drivers.

  • Capability: The system can detect unfaithful explanations. If a high-stakes agent (e.g., in medical or legal domains) provides a verbal justification that contradicts the reconstructed behavioral driver, the system triggers an immediate unfaithfulness alert, preventing users from relying on plausible but incorrect rationales.

2. Faithfulness-Constrained Reinforcement Learning (FCRL)

  • Improvement: Modify the fine-tuning objective (using DPO or RLHF) to include a Reasoning Alignment Loss. This loss function penalizes the divergence between the model's stated attribute importance and its actual choice sensitivity as measured by targeted occlusion (drop/equalize interventions).

  • Capability: The system produces Transparently Reasoned Agents. These models are mathematically incentivized to ensure that their verbalized logic is not merely a post-hoc imitation of rationality, but a true reflection of the probabilistic weights driving their decision-making process.

3. Transition from Natural Language Justification to Structured Argumentative Scoring

  • Improvement: Replace open-ended Direct Report prompting with a mandatory Score-Based Judge output protocol. The system must provide a structured vector of decisiveness scores (tau) for every attribute in a decision profile, rather than just a single sentence.

  • Capability: The system provides Audit-Ready Decision Logs. Instead of ambiguous text, the AI generates reproducible, quantitative argumentation frameworks that can be ingested by formal verification tools and human auditors to mathematically validate the decision logic.

4. Automated Reliability Certification via Targeted Occlusion Testing

  • Improvement: Embed an automated testing suite into the CI/CD pipeline that utilizes targeted occlusion (dropping or equalizing specific attributes) to stress-test the model's decision stability.

  • Capability: The system provides Sensitivity-Validated Deployment. Before any model update is pushed to production, it must pass a certification where its choice-flip rates under attribute perturbation match its stated priority rankings, ensuring that its behavior is consistent with its declared values.

Abstract

We ask whether large language models (LLMs) merely imitate rationales when choosing between two options, or whether their choices reflect a systematic underlying decision structure. Using synthetic binary decision settings in which models choose between profiles defined by graded attributes, we compare the attribute a model says mattered most with the attribute that best explains its choice under a behavioural model fit to prior decisions. The behavioural model predicts held-out choices well, showing that model behaviour is systematically related to the visible attributes rather than being random. However, direct self-reports and a separate score-based judge recover the behaviourally inferred driver only partially. The resulting picture is neither one of arbitrary behaviour nor one of fully articulated belief - outputs are structured enough to support prediction, but explicit reasons track the recovered driver only imperfectly. This qualitative pattern persists across prompt-order and sampling perturbations, alternative behavioural models, targeted occlusion analyses, and structurally varied decision settings. We interpret this as evidence for ``superficial belief'' in LLM decision-making: models behave as if guided by probabilistic local priorities over attributes, while having only limited verbal access to the attributes that drive their decisions.

Sources

Related papers