Out-of-Distribution Generalization of Risk Aversion in Language Models
cs.LG, cs.AI
Submitted: 2026-07-02
Updated: 2026-09-20
Comments: ICML 2026 Agents in the Wild workshop
Code: https://github.com/riskaverseAIs/riskaverseAIs
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
The gist: Training AIs to be risk-averse in resources could offer a failsafe in the event that AIs turn out misaligned.
Terminology
Abstract
Training AIs to be risk-averse in resources could offer a failsafe in the event that AIs turn out misaligned. Misaligned but risk-averse AIs would tend to prefer low-risk, low-reward strategies like cooperation over high-risk, high-reward strategies like rebellion, limiting the downsides of any misalignment. But we can only feasibly train AIs to be risk-averse on low-stakes gambles, and we will only be safe if their risk aversion generalizes to astronomically-high-stakes gambles. Will it? To shed light on this question, we introduce RiskAverseOOD: a benchmark for measuring how well risk aversion generalizes out of distribution. We then offer some initial results. Using a variety of methods to make Qwen3-8B choose risk-aversely when the stakes are low, we find that we can induce substantial risk aversion when the stakes are astronomically high. From a baseline 2% rate of choosing a safe `Cooperate' option, we see rates around 70% (SFT and tie training) and 52% (DPO). Activation steering scores 78% but hurts capabilities and makes the model excessively risk-averse. We observe similar effects at different scales (Qwen3-1.7B and Qwen3-14B) and across model families (Gemma-3-12B-IT and Llama-3.1-8B-Instruct). Overall, we find that risk aversion learned at low stakes can generalize OOD to astronomically high stakes, though not yet consistently enough to serve as a reliable failsafe. Achieving that level of consistency is an open problem.
Sources
- SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training
- Exploring Length Generalization in Large Language Models
- Tell me about yourself: LLMs are aware of their learned behaviors
- Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs
- LoRA: Low-Rank Adaptation of Large Language Models
- Natural Emergent Misalignment from Reward Hacking in Production RL
- Steering Out-of-Distribution Generalization with Concept Ablation Fine-Tuning
- Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs
- Persona Features Control Emergent Misalignment
- Spurious Correlation Learning in Preference Optimization: Mechanisms, Consequences, and Mitigation via Tie Training
- Risk Profiling and Modulation for LLMs
- Training language models to follow instructions with human feedback
- Steering Llama 2 via Contrastive Activation Addition
- Qwen3 Technical Report
- Direct Preference Optimization: Your Language Model is Secretly a Reward Model
- LLM economicus? Mapping the Behavioral Biases of LLMs via Utility Theory
- Model Organisms for Emergent Misalignment
- Fine-Tuning Language Models from Human Preferences
- Representation Engineering: A Top-Down Approach to AI Transparency
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks