Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning
Xujun Che, Yuchen Yuan, Weida Zhao, Chenyang Yu
cs.LG, cs.CL
Submitted: 2026-07-31
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Mitigating LLM Hallucinations via Conformal Abstention
- Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs
- Deep reinforcement learning from human preferences
- The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models
- Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty
- KARL: Mitigating Hallucinations in LLMs via Knowledge-Boundary-Aware Reinforcement Learning
- Scaling Laws for Reward Model Overoptimization
- Selective Classification for Deep Neural Networks
- A Theory of Regularized Markov Decision Processes
- On Calibration of Modern Neural Networks
- REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization
- AA-Omniscience: Evaluating Cross-Domain Knowledge Reliability in Large Language Models
- TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension
- Language Models (Mostly) Know What They Know
- Why Language Models Hallucinate
- Calibrated Language Models Must Hallucinate
- Selective Question Answering under Domain Shift
- Understanding the Effects of RLHF on LLM Generalisation and Diversity
- Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language Generation
- Reported Confidence in LLMs Tracks Commitment More Than Correctness
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks