Rethinking Prospect Theory for LLMs: Revealing the Instability of Decision-Making under Epistemic Uncertainty
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Rethinking Prospect Theory for LLMs".
Jane: Real-world decision-making often involves uncertainty expressed in linguistic rather than numerical terms, and Prospect Theory (PT) provides a classic framework for modeling human behavior under such uncertainty.
Tom: First, who's behind it and why it matters.
Title and authors: Tom: We started by looking at the title and the authors of "Rethinking Prospect Theory for LLMs: Revealing the Instability of Decision-Making under Epistemic Uncertainty." It tells us immediately that they aren't just applying existing theory; they are actively questioning if Prospect Theory is even a good lens for Large Language Model decision-making.
Jane: The authors are quite a team from institutions like Hong Kong University of Science and Technology, Huazhong University of Science and Technology, and the University of Illinois Urbana-Champaign, which tells us this work has a strong academic foundation.
Lu: I see the core idea is that existing research has tried to fit PT parameters to LLMs, but this paper zeroes in on whether PT actually describes the behavior when uncertainty comes in the form of vague language.
Meng: So, if you look at what they're trying to achieve, it seems their main goal is to show that these traditional PT frameworks can break down when the input is not a clean numerical probability.
Lalam: It really highlights that LLMs aren't just pattern matchers; they are making decisions based on an interpretation of language, and that interpretation isn't always stable.
The paper's summary: Tom: So, let’s talk about what the paper actually summarizes. They developed a three-stage workflow starting by estimating PT parameters from binary choices with precise probabilities to see how well the model captures behavior initially.
Jane: Then they moved into a second stage where they derived probability mappings for epistemic markers and injected those markers into prompts to test if the PT parameters stay stable under that linguistic uncertainty.
Lu: That mapping process, using their "No. Epistemic Marker Probability Mapping by Human," is crucial because it helps define how humans themselves map verbal expressions onto numerical probabilities before testing the model against that human baseline.
Meng: So they are essentially creating a controlled experiment where they test the limits of PT parameters when the input language gets messy, which makes sense for understanding deployment risks.
Lalam: It’s about seeing if those internal psychological structures of risk preference and loss aversion hold up when the uncertainty isn't a simple number but a word like "maybe."
The paper's improvements: Tom: Now, regarding the improvements they suggest for this work, they are clearly advocating for moving beyond just observing instability to actually developing ways to handle it better in real systems.
Jane: They emphasize that the current situation reveals a "representational non-invariance," meaning models alter their revealed preferences when equivalent uncertainty is expressed differently through linguistic markers.
Lu: The paper points out that while risk preference might stay relatively stable, the loss aversion and probability weighting shift more substantially under these conditions, which gives us a specific area to focus on for future work.
Meng: From a practical standpoint, this suggests that we need better ways to calibrate these models dynamically so they don't have such drastic shifts when facing natural language prompts in production.
Lalam: I think the key improvement they are pointing toward is needing a mechanism that can map linguistic markers directly into standardized numerical probabilities before applying any PT logic.
Conclusion: Tom: So, to wrap up, the core implication of "Rethinking Prospect Theory for LLMs: Revealing the Instability of Decision-Making under Epistemic Uncertainty" is that we can't just assume a model’s decision structure will remain consistent when it encounters ambiguous language.
Jane: The paper shows that for Prospect Theory to be a dependable tool, its validity has to hold up even when uncertainty is expressed through epistemic markers, which means models aren't universally adopting PT-like behavior.
Lu: I agree; the finding that large models exhibit more PT-like behavior but still have unstable loss aversion under linguistic uncertainty shows us exactly where their current limitations lie in interpreting human risk communication.
Meng: For implementation, this means we should be very careful about deploying these frameworks in safety-critical areas unless we can guarantee that the input is strictly translated into numeric values first.
Lalam: Ultimately, the paper reminds us that for safety-critical applications where epistemic ambiguity is everywhere, we have to stop treating verbal uncertainty as a direct substitute for numerical certainty and start building systems that handle it differently.
Hong Kong University of Science and Technology · Huazhong University of Science and Technology · University of Illinois Urbana-Champaign
cs.AI
Submitted: 2025-08-12
Updated: 2026-09-28
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 73/100
The gist: Real-world decision-making often involves uncertainty expressed in linguistic rather than numerical terms, and Prospect Theory (PT) provides a classic framework for modeling human behavior under such
Key concepts
- Prospect Theory (PT)
- A behavioral economics framework used to model how people make choices when facing risks and uncertainties. It describes how individuals value gains and losses differently, particularly emphasizing loss aversion, which means the pain of a loss is greater than the pleasure of an equivalent gain.
- Epistemic Uncertainty
- Uncertainty that arises from a lack of knowledge or information about the true state of affairs. In this study, it refers to situations where uncertainty is expressed through vague language or 'epistemic markers' instead of precise numerical probabilities, challenging how models interpret this ambiguity.
- Parameter Instability
- The finding that the specific parameters within Prospect Theory (like risk preference or loss aversion) change significantly when the input uncertainty is switched from numerical values to linguistic markers. This demonstrates that LLMs do not maintain consistent decision-making rules when faced with ambiguous language.
- Scale Dependency
- The observation that Prospect Theory-like behavior in LLMs only emerges reliably once the model reaches a sufficiently large parameter scale. Smaller models do not exhibit these behaviors consistently, indicating that complexity is necessary for PT adherence.
Terminology
Summary
Real-world decision-making often involves uncertainty expressed in linguistic rather than numerical terms, and Prospect Theory (PT) provides a classic framework for modeling human behavior under such uncertainty. This study develops a streamlined workflow to evaluate how Prospect Theory adequately describes Large Language Model (LLM) decision-making behavior and reveals the instability of these parameters when LLMs are exposed to epistemic uncertainty. The findings caution against deploying PT-based frameworks in real-world applications where epistemic ambiguity is prevalent, suggesting that LLMs do not exhibit universal Prospect-Theory-like behavior and their risk tendencies are likely not robust under linguistic uncertainty.
The core methodology involves a three-stage experiment designed to test the robustness of PT parameters against linguistic uncertainty.
-
First, the researchers
estimate PT parameters and evaluate how well the resulting model captures LLM decision-making behavior
by fitting PT parameters from binary choices with precise probabilities. -
Second, they
derive probability mappings for epistemic markers in the same context and inject them into prompts to examine the stability of PT parameters under linguistic uncertainty.
This stage involves having LLMs make binary choices between numerical probabilities and epistemic markers toinfer the model’s implicit probabilistic interpretation of those uncertainty expressions under the same context.
-
Finally, they
re-assess model behavior on the original decision tasks, now framed using epistemic markers grounded in their inferred probability values, to directly evaluate the impact of epistemic uncertainty on decision-making.
The framework is built upon classic behavioral economics to establish a reliable baseline for PT parameter measurement.
(Note: The paper details that they adopt a three-series lottery-choice experiment developed from Tanaka et al. (2010) to get a reliable PT parameter measurement,
where Series 1 and 2 elicit risk preference and probability weighting, while Series 3 is designed for loss aversion.)
The evaluation metrics used to assess model fit include:
(Note: The researchers quantify goodness-of-fit using McFadden pseudoR2 (McFadden, 1977), defined as R2 McFadden = 1 − LPT/Lnull,
and the mean absolute error (MAE) between the actual probability pactual and the predicted probability ppred of choosing option K.
)
The study investigates how epistemic markers influence LLM decision-making across varying degrees of linguistic uncertainty.
The researchers establish a No. Epistemic Marker Probability Mapping by Human
to define how humans map verbal expressions to numerical probabilities, which serves as the basis for subsequent experiments. They then perform Probability interpolation
using a slope formula to estimate an inferred probability mapping (pmapping) for each marker based on where the model reaches an implied probability of 0.5 (p0). This allows them to obtain a list of 14 probability values (one per marker) for each model, capturing how they semantically interpret verbal uncertainty in economic terms.
The results reveal key insights regarding the relationship between LLM scale, PT adherence, and epistemic uncertainty.
(Note: The findings show that Prospect Theory does not consistently perform well at explaining LLM decision behaviors,
and that larger models exhibit more Prospect-Theorylike behavior,
while loss aversion remains unstable.)
The study highlights several limitations:
-
Scale Dependency: PT-aligned behavior is
not an inherent model capability, but rather emerges only in models with sufficient parameter scale.
-
Representational Divergence: LLMs
exhibit severe cross-model divergence in absolute probability mappings,
even when capturing stable ordinal semantics across models. -
Parameter Instability: The introduction of linguistic uncertainty
profoundly destabilizes LLMs, rendering them unable to maintain theoretically coherent, stable, or interpretable PT parameters,
revealingrepresentational non-invariance.
The instability is further detailed by examining model-specific effects on PT fitness across four experimental rounds.
The analysis shows that while risk preference remains relatively stable despite marker substitution, loss aversion and probability weighting shift more substantially under these conditions.
Specifically, changes in λ suggest models do not preserve the same loss sensitivity when equivalent uncertainty is expressed through linguistic markers rather than numerical probabilities. Furthermore, drift under semantically matched substitutions reveals non-invariant decision-making: models alter revealed preferences when equivalent uncertainty is expressed differently.
The conclusion is that epistemic markers are not merely interchangeable verbal forms of numerical probabilities, but can directly affect the internal preference structure revealed by LLMs.
The study concludes with a strong recommendation for future system design and alignment.
The paper cautions against the uncritical deployment of PT-based frameworks to model or predict LLM behaviors in real-world applications where epistemic ambiguity is ubiquitous.
It suggests that for safety-critical applications, "epistemic uncertainty must be strictly translated into standardized numeric probabilities, or models must be explicitly aligned via persona-prompting to enforce desired risk profiles.
Improvements for AI systems
Here are specific improvements for AI systems based on the findings of this research, categorized by technical capability:
)1. Enhanced Robustness Against Linguistic Ambiguity (Epistemic Uncertainty)
LLMs should be equipped with a mechanism that recognizes and explicitly quantifies epistemic markers (e.g., likely,
somewhat unlikely
) as distinct from numerical probabilities during decision-making processes, rather than treating them as direct substitutes for numbers.
)2. Dynamic Parameter Calibration for Decision-Making
The system should implement a dynamic calibration layer that adjusts its internal Prospect Theory parameters (Risk Preference, Loss Aversion, and Probability Weighting) in real-time based on the linguistic context of the prompt. This prevents the observed Parameter Instability
where equivalent uncertainty expressed differently causes drastic shifts in preference structure.
)3. Scale-Dependent Confidence Thresholds
For safety-critical or high-stakes decisions (e.g., financial risk assessment, medical triage), the system should utilize a confidence threshold that is explicitly scaled with model size and robustness metrics (like McFadden R2 and MAE). Decisions based on PT frameworks should be flagged as Unreliable
if the model falls below a predetermined scale/robustness benchmark, preventing deployment of models exhibiting unstable decision-making.
)4. Contextual Uncertainty Translation Module
Develop a dedicated module that maps linguistic uncertainty markers to normalized probability distributions (using the derived model-specific mappings). This module should be used as an intermediary step: instead of feeding raw text prompts into the LLM's decision engine, the system first translates the epistemic language into a standardized numerical probability space before applying PT logic.
)5. Explicit Persona/Constraint Enforcement for Alignment
For applications requiring guaranteed risk profiles (e.g., regulatory compliance), systems must utilize explicit persona prompting or constraint rules to force models toward desired risk attitudes (e.g., forcing high loss aversion in financial contexts). This addresses the finding that native LLMs lack stable PT adherence without external conditioning.
)Improved AI System Capabilities:
The improved system will transition from a model that simply guesses
based on statistical artifacts to a system capable of:
-
Predicting how its own decision-making parameters will shift when faced with ambiguous, natural language inputs.
-
Maintaining decision stability across varying linguistic representations of uncertainty (e.g., handling both
highly likely
andprobable
without catastrophic shifts). -
Self-assessing the reliability of its own risk estimation framework before making a high-stakes choice, effectively acting as a meta-cognitive layer that warns when its underlying decision model is breaking down due to linguistic noise.
Sources
- Perceptions of Linguistic Uncertainty by Language Models and Humans
- Mathematical Proof as a Litmus Test: Revealing Failure Modes of Advanced Large Reasoning Models
- DeFine: Decision-Making with Analogical Reasoning over Factor Profiles
- Decision-Making Behavior Evaluation Framework for LLMs under Uncertain Context
- Are LLM-Judges Robust to Expressions of Uncertainty? Investigating the effect of Epistemic Markers on LLM-based Evaluation
- Learning to Ask Like a Physician
- A Conceptual Framework for AI-based Decision Systems in Critical Infrastructures
- Evaluating and Aligning Human Economic Risk Preferences in LLMs
- CostBench: Evaluating Multi-Turn Cost-Optimal Planning and Adaptation in Dynamic Environments for LLM Tool-Use Agents
- NOVA: NOise-aware Verbal Confidence CAlibration for Robust Large Language Models in RAG Systems
- Revisiting Epistemic Markers in Confidence Estimation: Can Markers Accurately Reflect Large Language Models' Uncertainty?
- An analysis of AI Decision under Risk: Prospect theory emerges in Large Language Models
- Qwen2.5 Technical Report
- Risk Profiling and Modulation for LLMs
- Diversity-Enhanced Reasoning for Subjective Questions
- Calibrating Behavioral Parameters with Large Language Models
- From Automation to Autonomy: A Survey on Large Language Models in Scientific Discovery
- Relying on the Unreliable: The Impact of Language Models' Reluctance to Express Uncertainty
- Navigating the Grey Area: How Expressions of Uncertainty and Overconfidence Affect Language Models
- CritiCal: Can Critique Help LLM Uncertainty or Confidence Calibration?
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection