Reward Valuation in Large Language Models: Causal Induction of Anhedonia
summary
The gist
Reward valuation in vision-language models is explored through a mechanistic framework inspired by clinical tests for anhedonia, revealing that perturbing specific reward-anticipatory units can
In short
Researchers found that specific reward-anticipatory units within vision-language models, analogous to human brain circuits, can be perturbed to induce behavioral effects mirroring anhedonia. This manipulation causes the AI to shift toward low-effort options when rewards are present but maintains general task capability otherwise. This establishes a causal link between internal reward representations and motivational deficits.
Key concepts
- Reward-Anticipatory Units
- These are specific neurons or circuits within the VLM that become highly active when the model predicts or anticipates receiving a reward. The study identifies these units by tracking activation changes during tasks with and without rewards, suggesting they are central to internal motivational states.
- Monetary Incentive Delay (MID) Paradigm
- This is a testing method used to isolate the internal motivational state of an AI. It measures the difference between when a reward is expected and when it is actually received. This helps researchers functionally localize which parts of the model are responsible for anticipating future rewards.
- Anhedonia-like Phenotype
- This refers to behavioral changes in the AI that resemble human anhedonia, specifically a shift toward choosing lower-effort, less rewarding options when incentives are available. Crucially, this shift happens without losing the model's overall ability to perform the task correctly under different conditions.
Terminology used across episodes
This episode discusses
- Reward Valuation in Large Language Models: Causal Induction of Anhedonia · Paper Radio
- The LLM Language Network: A Neuroscientific Approach for Identifying Causally Task-Relevant Units
- Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
- Scaling Laws for Task-Optimized Models of the Primate Visual Ventral Stream
- Measuring Massive Multitask Language Understanding
- Inducing Dyslexia in Vision Language Models · Paper Radio
- Depression Diagnosis Dialogue Simulation: Self-improving Psychiatrist with Tertiary Memory
- Do Large Language Models Think Like the Brain? Sentence-Level Evidences from Layer-Wise Embeddings and fMRI
- Contour Integration Underlies Human-Like Vision
- Instruction-tuning Aligns LLMs to the Human Brain
- A Diverse Corpus for Evaluating and Developing English Math Word Problem Solvers
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- TopoLM: brain-like spatio-functional organization in a topographic language model
- Alignment between Brains and AI: Evidence for Convergent Evolution across Modalities, Scales and Training Trajectories
- Interpreting and improving natural-language processing (in machines) with natural language-processing (in the brain)
- An Empirical Study of Example Forgetting during Deep Neural Network Learning
- Causal Mediation Analysis for Interpreting Neural NLP: The Case of Gender Bias
- PsychAdapter: Adapting LLM Transformers to Reflect Traits, Personality and Mental Health
- Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
The paper
Reward Valuation in Large Language Models: Causal Induction of Anhedonia · Read on arXiv
Melika Honarmand, Samin Mahdipour Aghabagher, Martin Schrimpf
NeuroAI Laboratory, EPFL
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "Reward Valuation in Large Language Models".
Tom: Reward valuation in vision-language models is explored through a mechanistic framework inspired by clinical tests for anhedonia, revealing that perturbing specific reward-anticipatory units can induce behavioral effects mirroring human anhedonia.
Jane: First, who's behind it and why it matters.
Title and authors: Tom: We’re starting out with the title of this paper, "Reward Valuation in Large Language Models: Causal Induction of Anhedonia," by Honarmand, Aghabagher, and Schrimpf, and it’s immediately clear they aren't just looking at how well the models perform tasks; they are trying to figure out the actual internal mechanism behind reward valuation.
Jane: That title is really telling because it moves beyond simple performance metrics and gets into the causal links between internal representations—how the AI values things—and a specific human experience, anhedonia. It’s about establishing a functional connection between computation and motivation.
Lu: The authors are drawing inspiration from clinical tests used to evaluate anhedonia in major depressive disorder, which is smart because it grounds their abstract AI analysis in established human psychology and neuroscience regarding the Nucleus Accumbens or NAc.
Meng: So, essentially, they’re taking these brain concepts—like how the NAc handles expected reward value during anticipation—and trying to find a functional parallel inside a large language model structure. That’s a big conceptual leap for an engineering team to tackle.
Lalam: It's exciting because it suggests that we might be able to map human motivational deficits onto AI structures, which opens up new avenues for designing systems with more nuanced internal states rather than just purely predictive ones.
The paper's summary: Tom: Now, summarizing the main part of "Reward Valuation in Large Language Models: Causal Induction of Anhedonia," the core finding is that they functionally identify reward-anticipatory units within vision-language models and show that perturbing them causes behavioral effects mirroring human anhedonia.
Jane: That means when they targeted those specific units, the model started acting like it was experiencing a motivational deficit, specifically by shifting toward less effortful or lower reward options during decision-making tasks.
Lu: The summary points out that this effect is not just some random glitch; it’s linked to how the NAc selectively encodes expected positive incentive value when anticipating a reward, and blunted activation in that area contributes directly to the clinical experience of anhedonia
Knutson et al., two thousand one Smids, two thousand twenty-three Daniels et al., two thousand twenty-five: .
Meng: So they are using activation differences during tasks with and without rewards to flag these sensitive units, specifically looking for those showing a significant increase in activation when they see reward-predicting stimuli. That’s the technical meat of the identification process.
Lalam: This summary is key because it proves that we can induce a specific motivational deficit purely by manipulating the AI's internal representations related to reward anticipation, which is a powerful demonstration of causality in this context.
The paper's improvements: Tom: What’s really interesting about the proposed improvements in "Reward Valuation in Large Language Models: Causal Induction of Anhedonia" is how they move beyond just showing an effect; they suggest that these identified NAc-selective units are essential for maintaining baseline motivation and drive.
Jane: They found that these specific units contribute a substantial amount to incentive direction, stating that they constitute only about zero point seven percent of all neurons in those targeted layers but account for six point nine percent of incentive direction, which is nearly ten times more than the uniform baseline activity.
Lu: That suggests these units act like a crucial volume knob for motivation; decreasing their scaling factors leads to a steady collapse in motivation, which is an important mechanistic detail for us to consider when designing future models.
Meng: From an engineering standpoint, knowing that these specific neurons are so critical gives us a target area for efficiency improvements; if we can maintain the function of those zero point seven percent efficiently while reducing overall model size, that’s a practical goal.
Lalam: The paper suggests we can build systems where this motivational structure is intentionally modulated, allowing us to control how much reward-seeking behavior an AI exhibits, which is a huge step toward controllability in complex agents.
Conclusion: Tom: So to wrap up "Reward Valuation in Large Language Models: Causal Induction of Anhedonia," the paper establishes a direct causal link between internal reward representations and motivational deficits, showing that perturbing those units induces anhedonia-like behavior in vision-language models.
Jane: The major implication is that this framework gives us a way to investigate psychiatric mechanisms through system-level perturbations in artificial neural networks, providing a computational basis for understanding why some systems might exhibit motivational failures.
Lu: This work really pushes the research toward developing interventions, as it identifies reliable markers for dysfunctional reward expectation that we can use to develop better criteria for those conditions
Arrondo et al., two thousand fifteen: .
Meng: For practical application, the study validates its findings through controls; they showed that when models were asked to calculate expected value explicitly, their conceptual understanding of reward utility stayed intact, which is a vital check against just observing a behavioral shift.
Lalam: It’s exciting because this work provides a foundation for exploring how these internal circuits translate into observable behaviors in AI agents, giving us new hypotheses for designing more robust and motivated learning systems.
More episodes
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language
- 2508.08833-An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems
- 2405.04118-Policy Learning with a Language Bottleneck