Risk-Conditioned Fine-Tuning of Large Language Models
cs.LG
Submitted: 2026-09-08
Updated: 2026-09-22
Comments: EMNLP 2026 Main
Code: https://github.com/SapanaChaudhary/RA-RLHF
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large Language Models (LLMs) are increasingly deployed in settings where rare but severe harmful generations can have significant consequences.
Terminology
Abstract
Large Language Models (LLMs) are increasingly deployed in settings where rare but severe harmful generations can have significant consequences. Existing Risk-Averse RLHF addresses this issue by optimizing Conditional Value-at-Risk (CVaR), but it trains policies for fixed risk levels and therefore cannot adjust the desired degree of risk aversion at inference time. In this paper, we propose risk-conditioned RLHF, a framework that trains a single policy that provides a continuous risk-control interface, enabling users to select different degrees of risk aversion without retraining or deploying multiple risk-specific models. Experiments across multiple benchmarks demonstrate that a single risk-conditioned policy can adapt to different risk levels at inference time, enabling more flexible and risk-aware LLM deployment.
Sources
- Expected Shortfall as a Tool for Financial Risk Management
- Portfolio Optimization with Spectral Measures of Risk
- PaLM 2 Technical Report
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
- Constitutional AI: Harmlessness from AI Feedback
- Evaluating Large Language Models Trained on Code
- Stochastic subgradient method converges at the rate $O(k^{-1/4})$ on weakly convex functions
- From Prompts to Protection: Large Language Model-Enabled In-Context Learning for Smart Public Safety UAV
- Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned
- RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models
- DisasterResponseGPT: Large Language Models for Accelerated Plan of Action Development in Disaster Response Scenarios
- The Llama 3 Herd of Models
- Personalized Soups: Personalized Large Language Model Alignment via Post-hoc Parameter Merging
- Continuous Language Model Interpolation for Dynamic and Controllable Text Generation
- Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes
- Multi-objective Large Language Model Alignment with Hierarchical Experts
- DeepSeek-V3 Technical Report
- Decoding-time Realignment of Language Models
- Enhancing LLM Safety via Constrained Direct Preference Optimization
- Robust General Utility for Reinforcement Learning
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks