Risk-Conditioned Fine-Tuning of Large Language Models

arXiv:2609.08064 · cs.LG · Submitted 2026-09-08 · Read on arXiv

cs.LG

Submitted: 2026-09-08

Updated: 2026-09-22

Comments: EMNLP 2026 Main

Code: https://github.com/SapanaChaudhary/RA-RLHF

License: http://creativecommons.org/licenses/by/4.0/

The gist: Large Language Models (LLMs) are increasingly deployed in settings where rare but severe harmful generations can have significant consequences.

Terminology

Abstract

Large Language Models (LLMs) are increasingly deployed in settings where rare but severe harmful generations can have significant consequences. Existing Risk-Averse RLHF addresses this issue by optimizing Conditional Value-at-Risk (CVaR), but it trains policies for fixed risk levels and therefore cannot adjust the desired degree of risk aversion at inference time. In this paper, we propose risk-conditioned RLHF, a framework that trains a single policy that provides a continuous risk-control interface, enabling users to select different degrees of risk aversion without retraining or deploying multiple risk-specific models. Experiments across multiple benchmarks demonstrate that a single risk-conditioned policy can adapt to different risk levels at inference time, enabling more flexible and risk-aware LLM deployment.

Sources

Related papers