BaNEL: Exploration Posteriors for Generative Modeling Using Only Negative Rewards
cs.LG, cs.AI
Submitted: 2025-10-10
Updated: 2026-09-08
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
The gist: Today's generative models thrive with large amounts of supervised data and informative reward functions characterizing the quality of the generation.
Terminology
Abstract
Today's generative models thrive with large amounts of supervised data and informative reward functions characterizing the quality of the generation. They work under the assumptions that the supervised data provides knowledge to pre-train the model, and the reward function provides dense information about how to further improve the generation quality and correctness. However, in the hardest instances of important problems, two problems arise: (1) the base generative model attains a near-zero reward signal, and (2) calls to the reward oracle are expensive. This setting poses a fundamentally different learning challenge than standard reward-based post-training. To address this, we propose BaNEL (Bayesian Negative Evidence Learning), an algorithm that post-trains the model using failed attempts only, while minimizing the number of reward evaluations (NREs). Our method is based on the idea that the problem of learning regularities underlying failures can be cast as another, in-loop generative modeling problem. We then leverage this model to assess whether new data resembles previously seen failures and steer the generation away from them. We show that BaNEL can improve model performance without observing a single successful sample on several sparse-reward tasks, outperforming existing novelty-bonus approaches by up to several orders of magnitude in success rate, while using fewer reward evaluations.
Sources
- Never Give Up: Learning Directed Exploration Strategies
- Flow Network based Generative Models for Non-Iterative Diverse Candidate Generation
- Exploration by Random Network Distillation
- Training Verifiers to Solve Math Word Problems
- RL$^2$: Fast Reinforcement Learning via Slow Reinforcement Learning
- Navigate the Unknown: Enhancing LLM Reasoning with Intrinsic Motivation Guided Exploration
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- $\pi$BO: Augmenting Acquisition Functions with User Beliefs for Bayesian Optimization
- Diffusion Model for Data-Driven Black-Box Optimization
- Hindsight policy gradients
- Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions
- On Entropy Control in LLM-RL Algorithms
- Incentivizing Exploration In Reinforcement Learning With Deep Predictive Models
- Qwen2.5 Technical Report
- Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning
- The Surprising Effectiveness of Negative Reinforcement in LLM Reasoning
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks