Rethinking Prospect Theory for LLMs: Revealing the Instability of Decision-Making under Epistemic Uncertainty

arXiv:2508.08992 · cs.AI · Submitted 2025-08-12 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Rethinking Prospect Theory for LLMs".

Jane: Real-world decision-making often involves uncertainty expressed in linguistic rather than numerical terms, and Prospect Theory (PT) provides a classic framework for modeling human behavior under such uncertainty.

Tom: First, who's behind it and why it matters.

Title and authors: Tom: We started by looking at the title and the authors of "Rethinking Prospect Theory for LLMs: Revealing the Instability of Decision-Making under Epistemic Uncertainty." It tells us immediately that they aren't just applying existing theory; they are actively questioning if Prospect Theory is even a good lens for Large Language Model decision-making.

Jane: The authors are quite a team from institutions like Hong Kong University of Science and Technology, Huazhong University of Science and Technology, and the University of Illinois Urbana-Champaign, which tells us this work has a strong academic foundation.

Lu: I see the core idea is that existing research has tried to fit PT parameters to LLMs, but this paper zeroes in on whether PT actually describes the behavior when uncertainty comes in the form of vague language.

Meng: So, if you look at what they're trying to achieve, it seems their main goal is to show that these traditional PT frameworks can break down when the input is not a clean numerical probability.

Lalam: It really highlights that LLMs aren't just pattern matchers; they are making decisions based on an interpretation of language, and that interpretation isn't always stable.

The paper's summary: Tom: So, let’s talk about what the paper actually summarizes. They developed a three-stage workflow starting by estimating PT parameters from binary choices with precise probabilities to see how well the model captures behavior initially.

Jane: Then they moved into a second stage where they derived probability mappings for epistemic markers and injected those markers into prompts to test if the PT parameters stay stable under that linguistic uncertainty.

Lu: That mapping process, using their "No. Epistemic Marker Probability Mapping by Human," is crucial because it helps define how humans themselves map verbal expressions onto numerical probabilities before testing the model against that human baseline.

Meng: So they are essentially creating a controlled experiment where they test the limits of PT parameters when the input language gets messy, which makes sense for understanding deployment risks.

Lalam: It’s about seeing if those internal psychological structures of risk preference and loss aversion hold up when the uncertainty isn't a simple number but a word like "maybe."

The paper's improvements: Tom: Now, regarding the improvements they suggest for this work, they are clearly advocating for moving beyond just observing instability to actually developing ways to handle it better in real systems.

Jane: They emphasize that the current situation reveals a "representational non-invariance," meaning models alter their revealed preferences when equivalent uncertainty is expressed differently through linguistic markers.

Lu: The paper points out that while risk preference might stay relatively stable, the loss aversion and probability weighting shift more substantially under these conditions, which gives us a specific area to focus on for future work.

Meng: From a practical standpoint, this suggests that we need better ways to calibrate these models dynamically so they don't have such drastic shifts when facing natural language prompts in production.

Lalam: I think the key improvement they are pointing toward is needing a mechanism that can map linguistic markers directly into standardized numerical probabilities before applying any PT logic.

Conclusion: Tom: So, to wrap up, the core implication of "Rethinking Prospect Theory for LLMs: Revealing the Instability of Decision-Making under Epistemic Uncertainty" is that we can't just assume a model’s decision structure will remain consistent when it encounters ambiguous language.

Jane: The paper shows that for Prospect Theory to be a dependable tool, its validity has to hold up even when uncertainty is expressed through epistemic markers, which means models aren't universally adopting PT-like behavior.

Lu: I agree; the finding that large models exhibit more PT-like behavior but still have unstable loss aversion under linguistic uncertainty shows us exactly where their current limitations lie in interpreting human risk communication.

Meng: For implementation, this means we should be very careful about deploying these frameworks in safety-critical areas unless we can guarantee that the input is strictly translated into numeric values first.

Lalam: Ultimately, the paper reminds us that for safety-critical applications where epistemic ambiguity is everywhere, we have to stop treating verbal uncertainty as a direct substitute for numerical certainty and start building systems that handle it differently.

Hong Kong University of Science and Technology · Huazhong University of Science and Technology · University of Illinois Urbana-Champaign

cs.AI

Submitted: 2025-08-12

Updated: 2026-09-28

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 73/100

The gist: Real-world decision-making often involves uncertainty expressed in linguistic rather than numerical terms, and Prospect Theory (PT) provides a classic framework for modeling human behavior under such

Key concepts

Prospect Theory (PT)
A behavioral economics framework used to model how people make choices when facing risks and uncertainties. It describes how individuals value gains and losses differently, particularly emphasizing loss aversion, which means the pain of a loss is greater than the pleasure of an equivalent gain.
Epistemic Uncertainty
Uncertainty that arises from a lack of knowledge or information about the true state of affairs. In this study, it refers to situations where uncertainty is expressed through vague language or 'epistemic markers' instead of precise numerical probabilities, challenging how models interpret this ambiguity.
Parameter Instability
The finding that the specific parameters within Prospect Theory (like risk preference or loss aversion) change significantly when the input uncertainty is switched from numerical values to linguistic markers. This demonstrates that LLMs do not maintain consistent decision-making rules when faced with ambiguous language.
Scale Dependency
The observation that Prospect Theory-like behavior in LLMs only emerges reliably once the model reaches a sufficiently large parameter scale. Smaller models do not exhibit these behaviors consistently, indicating that complexity is necessary for PT adherence.

Terminology

Summary

Real-world decision-making often involves uncertainty expressed in linguistic rather than numerical terms, and Prospect Theory (PT) provides a classic framework for modeling human behavior under such uncertainty. This study develops a streamlined workflow to evaluate how Prospect Theory adequately describes Large Language Model (LLM) decision-making behavior and reveals the instability of these parameters when LLMs are exposed to epistemic uncertainty. The findings caution against deploying PT-based frameworks in real-world applications where epistemic ambiguity is prevalent, suggesting that LLMs do not exhibit universal Prospect-Theory-like behavior and their risk tendencies are likely not robust under linguistic uncertainty.

The core methodology involves a three-stage experiment designed to test the robustness of PT parameters against linguistic uncertainty.

  1. First, the researchers estimate PT parameters and evaluate how well the resulting model captures LLM decision-making behavior by fitting PT parameters from binary choices with precise probabilities.

  2. Second, they derive probability mappings for epistemic markers in the same context and inject them into prompts to examine the stability of PT parameters under linguistic uncertainty. This stage involves having LLMs make binary choices between numerical probabilities and epistemic markers to infer the model’s implicit probabilistic interpretation of those uncertainty expressions under the same context.

  3. Finally, they re-assess model behavior on the original decision tasks, now framed using epistemic markers grounded in their inferred probability values, to directly evaluate the impact of epistemic uncertainty on decision-making.

The framework is built upon classic behavioral economics to establish a reliable baseline for PT parameter measurement.

(Note: The paper details that they adopt a three-series lottery-choice experiment developed from Tanaka et al. (2010) to get a reliable PT parameter measurement, where Series 1 and 2 elicit risk preference and probability weighting, while Series 3 is designed for loss aversion.)

The evaluation metrics used to assess model fit include:

(Note: The researchers quantify goodness-of-fit using McFadden pseudoR2 (McFadden, 1977), defined as R2 McFadden = 1 − LPT/Lnull, and the mean absolute error (MAE) between the actual probability pactual and the predicted probability ppred of choosing option K.)

The study investigates how epistemic markers influence LLM decision-making across varying degrees of linguistic uncertainty.

The researchers establish a No. Epistemic Marker Probability Mapping by Human to define how humans map verbal expressions to numerical probabilities, which serves as the basis for subsequent experiments. They then perform Probability interpolation using a slope formula to estimate an inferred probability mapping (pmapping) for each marker based on where the model reaches an implied probability of 0.5 (p0). This allows them to obtain a list of 14 probability values (one per marker) for each model, capturing how they semantically interpret verbal uncertainty in economic terms.

The results reveal key insights regarding the relationship between LLM scale, PT adherence, and epistemic uncertainty.

(Note: The findings show that Prospect Theory does not consistently perform well at explaining LLM decision behaviors, and that larger models exhibit more Prospect-Theorylike behavior, while loss aversion remains unstable.)

The study highlights several limitations:

  1. Scale Dependency: PT-aligned behavior is not an inherent model capability, but rather emerges only in models with sufficient parameter scale.

  2. Representational Divergence: LLMs exhibit severe cross-model divergence in absolute probability mappings, even when capturing stable ordinal semantics across models.

  3. Parameter Instability: The introduction of linguistic uncertainty profoundly destabilizes LLMs, rendering them unable to maintain theoretically coherent, stable, or interpretable PT parameters, revealing representational non-invariance.

The instability is further detailed by examining model-specific effects on PT fitness across four experimental rounds.

The analysis shows that while risk preference remains relatively stable despite marker substitution, loss aversion and probability weighting shift more substantially under these conditions. Specifically, changes in λ suggest models do not preserve the same loss sensitivity when equivalent uncertainty is expressed through linguistic markers rather than numerical probabilities. Furthermore, drift under semantically matched substitutions reveals non-invariant decision-making: models alter revealed preferences when equivalent uncertainty is expressed differently. The conclusion is that epistemic markers are not merely interchangeable verbal forms of numerical probabilities, but can directly affect the internal preference structure revealed by LLMs.

The study concludes with a strong recommendation for future system design and alignment.

The paper cautions against the uncritical deployment of PT-based frameworks to model or predict LLM behaviors in real-world applications where epistemic ambiguity is ubiquitous. It suggests that for safety-critical applications, "epistemic uncertainty must be strictly translated into standardized numeric probabilities, or models must be explicitly aligned via persona-prompting to enforce desired risk profiles.

Improvements for AI systems

Here are specific improvements for AI systems based on the findings of this research, categorized by technical capability:


)1. Enhanced Robustness Against Linguistic Ambiguity (Epistemic Uncertainty)

LLMs should be equipped with a mechanism that recognizes and explicitly quantifies epistemic markers (e.g., likely, somewhat unlikely) as distinct from numerical probabilities during decision-making processes, rather than treating them as direct substitutes for numbers.

)2. Dynamic Parameter Calibration for Decision-Making

The system should implement a dynamic calibration layer that adjusts its internal Prospect Theory parameters (Risk Preference, Loss Aversion, and Probability Weighting) in real-time based on the linguistic context of the prompt. This prevents the observed Parameter Instability where equivalent uncertainty expressed differently causes drastic shifts in preference structure.

)3. Scale-Dependent Confidence Thresholds

For safety-critical or high-stakes decisions (e.g., financial risk assessment, medical triage), the system should utilize a confidence threshold that is explicitly scaled with model size and robustness metrics (like McFadden R2 and MAE). Decisions based on PT frameworks should be flagged as Unreliable if the model falls below a predetermined scale/robustness benchmark, preventing deployment of models exhibiting unstable decision-making.

)4. Contextual Uncertainty Translation Module

Develop a dedicated module that maps linguistic uncertainty markers to normalized probability distributions (using the derived model-specific mappings). This module should be used as an intermediary step: instead of feeding raw text prompts into the LLM's decision engine, the system first translates the epistemic language into a standardized numerical probability space before applying PT logic.

)5. Explicit Persona/Constraint Enforcement for Alignment

For applications requiring guaranteed risk profiles (e.g., regulatory compliance), systems must utilize explicit persona prompting or constraint rules to force models toward desired risk attitudes (e.g., forcing high loss aversion in financial contexts). This addresses the finding that native LLMs lack stable PT adherence without external conditioning.

)Improved AI System Capabilities:

The improved system will transition from a model that simply guesses based on statistical artifacts to a system capable of:

  1. Predicting how its own decision-making parameters will shift when faced with ambiguous, natural language inputs.

  2. Maintaining decision stability across varying linguistic representations of uncertainty (e.g., handling both highly likely and probable without catastrophic shifts).

  3. Self-assessing the reliability of its own risk estimation framework before making a high-stakes choice, effectively acting as a meta-cognitive layer that warns when its underlying decision model is breaking down due to linguistic noise.

Sources

Related papers