Reward Valuation in Large Language Models: Causal Induction of Anhedonia

arXiv:2607.06626 · cs.LG, q-bio.NC · Submitted 2026-07-07 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "Reward Valuation in Large Language Models".

Tom: Reward valuation in vision-language models is explored through a mechanistic framework inspired by clinical tests for anhedonia, revealing that perturbing specific reward-anticipatory units can induce behavioral effects mirroring human anhedonia.

Jane: First, who's behind it and why it matters.

Title and authors: Tom: We’re starting out with the title of this paper, "Reward Valuation in Large Language Models: Causal Induction of Anhedonia," by Honarmand, Aghabagher, and Schrimpf, and it’s immediately clear they aren't just looking at how well the models perform tasks; they are trying to figure out the actual internal mechanism behind reward valuation.

Jane: That title is really telling because it moves beyond simple performance metrics and gets into the causal links between internal representations—how the AI values things—and a specific human experience, anhedonia. It’s about establishing a functional connection between computation and motivation.

Lu: The authors are drawing inspiration from clinical tests used to evaluate anhedonia in major depressive disorder, which is smart because it grounds their abstract AI analysis in established human psychology and neuroscience regarding the Nucleus Accumbens or NAc.

Meng: So, essentially, they’re taking these brain concepts—like how the NAc handles expected reward value during anticipation—and trying to find a functional parallel inside a large language model structure. That’s a big conceptual leap for an engineering team to tackle.

Lalam: It's exciting because it suggests that we might be able to map human motivational deficits onto AI structures, which opens up new avenues for designing systems with more nuanced internal states rather than just purely predictive ones.

The paper's summary: Tom: Now, summarizing the main part of "Reward Valuation in Large Language Models: Causal Induction of Anhedonia," the core finding is that they functionally identify reward-anticipatory units within vision-language models and show that perturbing them causes behavioral effects mirroring human anhedonia.

Jane: That means when they targeted those specific units, the model started acting like it was experiencing a motivational deficit, specifically by shifting toward less effortful or lower reward options during decision-making tasks.

Lu: The summary points out that this effect is not just some random glitch; it’s linked to how the NAc selectively encodes expected positive incentive value when anticipating a reward, and blunted activation in that area contributes directly to the clinical experience of anhedonia

Knutson et al., two thousand one Smids, two thousand twenty-three Daniels et al., two thousand twenty-five: .

Meng: So they are using activation differences during tasks with and without rewards to flag these sensitive units, specifically looking for those showing a significant increase in activation when they see reward-predicting stimuli. That’s the technical meat of the identification process.

Lalam: This summary is key because it proves that we can induce a specific motivational deficit purely by manipulating the AI's internal representations related to reward anticipation, which is a powerful demonstration of causality in this context.

The paper's improvements: Tom: What’s really interesting about the proposed improvements in "Reward Valuation in Large Language Models: Causal Induction of Anhedonia" is how they move beyond just showing an effect; they suggest that these identified NAc-selective units are essential for maintaining baseline motivation and drive.

Jane: They found that these specific units contribute a substantial amount to incentive direction, stating that they constitute only about zero point seven percent of all neurons in those targeted layers but account for six point nine percent of incentive direction, which is nearly ten times more than the uniform baseline activity.

Lu: That suggests these units act like a crucial volume knob for motivation; decreasing their scaling factors leads to a steady collapse in motivation, which is an important mechanistic detail for us to consider when designing future models.

Meng: From an engineering standpoint, knowing that these specific neurons are so critical gives us a target area for efficiency improvements; if we can maintain the function of those zero point seven percent efficiently while reducing overall model size, that’s a practical goal.

Lalam: The paper suggests we can build systems where this motivational structure is intentionally modulated, allowing us to control how much reward-seeking behavior an AI exhibits, which is a huge step toward controllability in complex agents.

Conclusion: Tom: So to wrap up "Reward Valuation in Large Language Models: Causal Induction of Anhedonia," the paper establishes a direct causal link between internal reward representations and motivational deficits, showing that perturbing those units induces anhedonia-like behavior in vision-language models.

Jane: The major implication is that this framework gives us a way to investigate psychiatric mechanisms through system-level perturbations in artificial neural networks, providing a computational basis for understanding why some systems might exhibit motivational failures.

Lu: This work really pushes the research toward developing interventions, as it identifies reliable markers for dysfunctional reward expectation that we can use to develop better criteria for those conditions

Arrondo et al., two thousand fifteen: .

Meng: For practical application, the study validates its findings through controls; they showed that when models were asked to calculate expected value explicitly, their conceptual understanding of reward utility stayed intact, which is a vital check against just observing a behavioral shift.

Lalam: It’s exciting because this work provides a foundation for exploring how these internal circuits translate into observable behaviors in AI agents, giving us new hypotheses for designing more robust and motivated learning systems.

Melika Honarmand, Samin Mahdipour Aghabagher, Martin Schrimpf

NeuroAI Laboratory, EPFL

cs.LG, q-bio.NC

Submitted: 2026-07-07

Updated: 2026-09-30

Code: https://github.com/epflneuroailab/Anhedonic-AI

License: http://creativecommons.org/licenses/by-nc-sa/4.0/

Importance score: 89/100

The gist: Reward valuation in vision-language models is explored through a mechanistic framework inspired by clinical tests for anhedonia, revealing that perturbing specific reward-anticipatory units can

Key concepts

Reward-Anticipatory Units
These are specific neurons or circuits within the VLM that become highly active when the model predicts or anticipates receiving a reward. The study identifies these units by tracking activation changes during tasks with and without rewards, suggesting they are central to internal motivational states.
Monetary Incentive Delay (MID) Paradigm
This is a testing method used to isolate the internal motivational state of an AI. It measures the difference between when a reward is expected and when it is actually received. This helps researchers functionally localize which parts of the model are responsible for anticipating future rewards.
Anhedonia-like Phenotype
This refers to behavioral changes in the AI that resemble human anhedonia, specifically a shift toward choosing lower-effort, less rewarding options when incentives are available. Crucially, this shift happens without losing the model's overall ability to perform the task correctly under different conditions.

Terminology

Summary

Reward valuation in vision-language models is explored through a mechanistic framework inspired by clinical tests for anhedonia, revealing that perturbing specific reward-anticipatory units can induce behavioral effects mirroring human anhedonia. This research establishes a causal link between internal reward representations in AI and motivational deficits, suggesting that these circuits parallel those found in the human brain's Nucleus Accumbens.

How it works

The researchers use neuroscientifically inspired methods to functionally identify reward-anticipatory circuits within vision-language models (VLMs) and evaluate their causal involvement via targeted perturbations. This approach is built on the idea that NAc deficits, linked to anhedonia in major depressive disorder, can be functionally analogous to specific units in AI models.

The methodology involves several key steps:

  1. Identification of reward-sensitive units by comparing model activations during tasks with and without reward incentives, specifically looking for units exhibiting a significant increase in activation (∆ > 3σ) in response to reward-predicting stimuli.

  2. Functional localization using the Monetary Incentive Delay (MID) paradigm, which separates the expectation of a reward from its eventual receipt to capture internal motivational states during an anticipatory delay phase.

  3. Evaluation of decision-making and psychometric profiles following activation patching of NAc-selective units to test their causal role in inducing behavioral effects.

Key Findings and Behavioral Manifestations

Perturbing the identified NAc-selective units induces behavioral effects that mirror human anhedonia: the model shifts toward low-effort, low-reward options in effort-based decision-making tasks. Crucially, the perturbed model maintains baseline performance when reward-based choice is removed, reflecting a specific deficit in reward valuation and anticipation rather than a loss of task capability. This induced vulnerability aligns with clinical anhedonia and motivation scales, including DARS and MAP-SR.

The behavioral shifts are quantified across several benchmarks:

(a) Clinical Psychometric Profile:

(b) Effort-Based Decision Making:

The ASDiv-EEfRT showed that the perturbed model demonstrated a significant preference shift toward the lower-effort tasks, resulting in a marked reduction in the mean points attempted. However, when tested under a forced-choice control, the perturbed model maintained its baseline accuracy, confirming that this shift is driven by reward-seeking motivation rather than general reasoning decline.

Mechanistic Insights into Circuit Function

The analysis reveals a functional layer dissociation within the VLM architecture:

  1. Mid-layer neurons (layers 13–14) exhibited a marked suppression, with activations plunging significantly below neutral baseline levels.

  2. Late-layer neurons (layers 18–27) demonstrated elevated activation in response to reward cues, suggesting these are the units targeted for primary anhedonia induction.

Furthermore, NAc-selective units act as the primary contributor to the incentive direction; they constitute only 0.7% of all neurons in targeted layers yet contribute 6.9% of incentive direction, 9.5 times more than the uniform baseline. This demonstrates that these specific units are essential for maintaining baseline motivation and drive, acting as a functional volume knob for motivation where decreasing scaling factors lead to a steady collapse in motivation.

Validation and Generalization

The study validates the findings through several controls:

(a) Control Experiments:

(b) Reward Knowledge and Perspective Shifting:

To ensure the behavioral shift resulted from a deficit in valuation rather than calculation, models were asked to explicitly calculate expected value (EV), where both intact and perturbed models correctly identified the HE/HR task as having the superior EV. When choosing for a third party, the perturbed model's selection rate remained largely consistent with intact model, indicating preserved conceptual understanding of reward utility.

(c) Domain Generalization:

The anhedonic behavior generalized across ten diverse MMLU domains, where the perturbed model consistently chooses lower-reward options regardless of the discipline. This effect is robust, as capability controls showed that accuracy remained stable in nine out of ten domains when no choice-reward element was present.

Conclusion

The work establishes that VLMs possess internal reward-anticipatory units that directly influence behavioral output in a manner aligned with human data. By applying targeted perturbations to these units, the researchers demonstrated a direct causal link between internal reward representations and the emergence of anhedonia-like phenotypes in silico, suggesting that these circuits are essential for maintaining goal-directed motivation. This framework provides a foundation for exploring psychiatric mechanisms through system-level perturbations in artificial neural networks.


The gist

Perturbing specific reward-anticipatory units in vision-language models induces behavioral effects mirroring human anhedonia by shifting the model toward low-effort, low-reward options while preserving general task capability.

Improvements for AI systems

Here are the specific, high-impact improvements that can be made to AI systems based on this research:


  1. Improvement: Implement a Reward-Anticipatory Unit (RAU) Module within Vision-Language Models (VLMs). This involves developing a functional localization technique (using activation delta thresholds, e.g., 3σ) to identify units in the model's transformer layers that mirror the Nucleus Accumbens (NAc) in human reward circuitry.

  2. Improvement: Develop a Motivational Perturbation capability. By selectively patching or perturbing these identified NAc-selective units with neutral baseline activations, the system can be induced to exhibit anhedonia-like behavioral deficits (reduced motivation and pleasure anticipation) without destroying its core task-solving capabilities (as shown by the control experiments).

  3. Improvement: Enhance Decision-Making Under Risk and Effort Trade-offs. The system should be trained to understand the cost-benefit imbalance between effort and reward magnitude, mirroring the clinical observation that individuals with anhedonia fail to expend effort for uncertain rewards. This is achieved by explicitly training models on paradigms like the Probability-EEfRT, where they must weigh a fixed low-effort/low-reward option against a high-effort/high-reward option under uncertainty.

  4. Improvement: Develop Robust Domain Generalization of Anhedonic Behavior. The motivational deficit should be generalized across diverse tasks (MMLU domains). The system should consistently show a preference for lower-reward options regardless of the subject matter (e.g., choosing a low-difficulty math problem over an easy global facts question when reward is scaled by difficulty), confirming that the lack of motivation is not task-specific but a systemic property.

  5. Improvement: Create Clinical Diagnostic and Intervention Proxies. The model can serve as a digital twin for in silico psychiatric research. By observing the specific pattern of NAc-selective unit suppression, researchers can establish a direct computational link between internal representations and observable motivational deficits (like those measured by DARS or MAP-SR), providing novel hypotheses for understanding MDD and anhedonia mechanisms.

The resulting improved AI system can:

  1. Perform complex reasoning while maintaining robust cognitive performance even when its reward-seeking motivation is selectively impaired.

  2. Make optimal, effort-aware decisions in uncertain environments by accurately calculating the trade-off between task difficulty and potential reward magnitude, rather than just maximizing immediate gain.

  3. Act as a safety and diagnostic tool in developing future AI agents for complex human interaction scenarios by identifying internal representations that correlate with motivational deficits relevant to conditions like depression or anhedonia.

Abstract

Recent frontier models mimic complex aspects of human cognition. Here we ask whether this alignment extends into reward valuation, which we assess in a mechanistic framework. Specifically, we use clinical tests that were developed to evaluate anhedonia in human subjects with major depressive disorders. Mechanistically, anhedonia is frequently associated with dysregulation in the Nucleus Accumbens (NAc) and the broader dopaminergic reward system. While neuroimaging has localized these deficits, establishing a causal link between NAc activity and specific behavioral symptoms remains a challenge. We use these ideas from neuroscience to functionally identify reward-anticipatory units in state-of-the-art AI models, and evaluate their causal involvement via targeted perturbations. We find that not only are such model units predictive of NAc brain recordings, their perturbation also induces behavioral effects mirroring human anhedonia: the model opts for low-effort, low-reward tasks in effort-based decision-making paradigms. Crucially, our results demonstrate that this represents a specific deficit in self-centered reward valuation and anticipation--rather than a loss of task capability, reward calculation, or effort avoidance. This induced vulnerability aligns with clinical measures of anhedonia and motivation in humans, such as DARS and MAP-SR, instruments that contain no reward-related vocabulary, ruling out a purely lexical account of the perturbation effect. Taken together, our results suggest reward valuation circuits in AI models that functionally mimic those in humans.

Sources

Related papers