Reinforcement Learning over Predictive Distributions for LLM Regression
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Reinforcement Learning over Predictive Distributions for LLM Regression".
Jane: Large language models can predict real-valued quantities from heterogeneous inputs such as text, code, and molecular strings, but most training objectives score each decoded floating-point number independently,
Tom: First, who's behind it and why it matters.
Title and authors: Tom: Alright, so we're starting with the title and who's behind this paper, "Reinforcement Learning over Predictive Distributions for LLM Regression." It tells us immediately that the core idea is using reinforcement learning specifically to shape how these models create their probability distributions for regression tasks.
Jane: And the authors are Jungsoo Park, Hyungjoo Chae, Ethan Mendes, Jay DeYoung, Varsha Kishore, and Wei Xu. They’re a solid group of researchers from places like Georgia Institute of Technology and Allen Institute for AI.
Lu: I find it interesting that they focused on regression tasks specifically because predicting real-valued quantities is where we really need calibrated uncertainty estimates to make sense of the results.
Meng: When you look at the abstract, they point out that most current training objectives just score each decoded floating-point number independently, which means they improve those point estimates but don't guarantee a calibrated predictive distribution, which is a big problem for applications requiring ranking or uncertainty estimation.
Lalam: It seems like this paper is directly addressing that limitation by introducing Distribution-Aware Reward to train the language models to produce better distributions instead of just optimizing individual outputs against scalar targets.
The paper's summary: Tom: So, looking at what they actually propose in the paper, Jungsoo Park et al. introduce Distribution-Aware Reward as an on-policy reinforcement learning objective designed to train language models to produce better predictive distributions for regression tasks rather than just optimizing individual decoded outputs against scalar targets.
Jane: In simpler terms, they are treating all the different predictions the model makes during training as a collection, or an empirical predictive distribution, and they evaluate that whole collection using something called the Continuous Ranked Probability Score, which is a proper way to score distributions.
Lu: That CRPS evaluation is key because it gives them a principled way to check how good their resulting distributions are at actually predicting the target values.
Meng: And they assign credit for each individual rollout using something called leave-one-out credit assignment, which means they figure out how much each specific prediction contributes to making the whole distribution better, rewarding predictions that improve the dispersion around the ground truth.
Lalam: It’s about moving from optimizing single guesses to optimizing the quality of the entire set of guesses, which is a significant shift in what we train these models for.
The paper's improvements: Tom: One of the main improvements they highlight is that pointwise rewards, like standard MSE, score each rollout independently and actually encourage predictions to collapse toward the mean, which can lead to predictions that are poorly calibrated or overly narrow.
Jane: But Distribution-Aware Reward does something different; it scores each rollout by its leave-one-out contribution to the full predictive distribution, which rewards both accuracy and useful spread around the ground truth.
Lu: They show that this results in rollouts that are both accurate and better calibrated as a set, which improves both pointwise error and rank correlation, especially when looking at how they perform on interpolation and extrapolation behavior in synthetic Gaussian-mixture tasks.
Meng: This is interesting because it suggests that optimizing the predictive distribution formed by multiple LLM rollouts leads to more reliable regression behavior than just optimizing individual decoded outputs against scalar targets.
Lalam: It means we’re getting predictions that are not just a better single number, but a set of numbers whose collective behavior is much more trustworthy for real-world use.
Conclusion: Tom: So to wrap up on "Reinforcement Learning over Predictive Distributions for LLM Regression," the authors demonstrate that DAR optimizes the predictive distribution formed by multiple LLM rollouts, moving past optimizing individual decoded outputs to match scalar targets.
Jane: They conclude that this method encourages predictions that are both accurate and better calibrated as a set, which improves both pointwise error and rank correlation across different settings like synthetic tasks and molecular property prediction.
Lu: The analysis of the resulting predictive distributions also shows that the rollout dispersion better reflects example-level prediction difficulty, indicating their spread is more meaningful than just a generic measure of uncertainty.
Meng: They also observed that under training with GRPO, DAR encourages more useful exploration by achieving a higher bracketing rate and maintaining higher rollout entropy during training, which helps mitigate diversity collapse.
Lalam: Ultimately, the paper shows that optimizing these predictive distributions provides more reliable regression behavior than pointwise objectives alone, offering a promising direction for using LLMs as flexible regressors in scientific fields.
Georgia Institute of Technology · Allen Institute for AI
cs.LG, cs.AI, cs.CL
Submitted: 2026-05-20
Updated: 2026-10-06
Code: https://github.com/verl-project/verl
Importance score: 90/100
The gist: Large language models can predict real-valued quantities from heterogeneous inputs such as text, code, and molecular strings, but most training objectives score each decoded floating-point number
Key concepts
- Continuous Ranked Probability Score (CRPS)
- CRPS is a proper scoring rule used to measure how good a predicted probability distribution is compared to the true target value. By maximizing the negative CRPS, DAR encourages the model's predictions to be not just accurate but also appropriately dispersed, ensuring it captures uncertainty well.
- Leave-One-Out Credit Assignment
- This technique assigns credit for a prediction by comparing the score of the full set of predictions against a set where one specific prediction is removed. This helps determine how much each individual rollout contributes to the overall quality of the predictive distribution.
- Pointwise Rewards vs. DAR
- Pointwise rewards, like Mean Squared Error (MSE), treat each output independently, often leading to narrow or poorly calibrated predictions. DAR contrasts this by optimizing the collective predictive distribution, encouraging models to be accurate and better calibrated as a set.
- Predictive Distribution
- Instead of focusing on a single decoded number, DAR treats all multiple LLM rollouts from the same input as an empirical predictive distribution. This allows the objective function to evaluate how well the model's predictions cover the target space, leading to better uncertainty estimates.
Terminology
Summary
Large language models can predict real-valued quantities from heterogeneous inputs such as text, code, and molecular strings, but most training objectives score each decoded floating-point number independently, improving point estimates without ensuring calibrated predictive distributions.
The gist
Distribution-Aware Reward (DAR) is an on-policy reinforcement learning objective that trains language models to produce better predictive distributions for regression tasks by treating multiple decoded samples as an empirical predictive distribution and evaluating it with the Continuous Ranked Probability Score (CRPS).
How it works
-
The method treats
multiple decoded samples as an empirical predictive distribution.
-
It evaluates this distribution against the target using the
Continuous Ranked Probability Score (CRPS),
which is defined as a proper scoring rule for predictive distributions. The objective is to maximize the negated CRPS, denoted asSi = −CRPS(Fi, yi).
-
To assign credit to individual rollouts, DAR uses
leave-one-out credit assignment.
For each rollout prediction, it calculates the marginal contribution by comparing the score of the full distribution with aleave-one-out empirical distribution
obtained by removing that specific rollout:R DAR ik = Si − S(-k) i.
-
This mechanism rewards rollouts that are both accurate and appropriately dispersed, encouraging predictions that are
accurate and better calibrated as a set,
rather than just optimizing individual outputs against scalar targets.
Key Contributions to Optimization
(Based on the comparison between pointwise rewards and DAR)
-
Pointwise rewards (e.g., MSE)
score each rollout independently, encouraging predictions to collapse toward the mean.
This can lead to predictions that arepoorly calibrated, overly narrow, or one-sided,
which weakens uncertainty estimation and rank-based evaluation. -
DAR optimizes the
predictive distribution formed by multiple LLM rollouts,
moving beyond optimizingindividual decoded outputs to match scalar targets.
-
DAR encourages predictions that are
accurate and better calibrated as a set (Figure 1), improving both pointwise error and rank correlation.
Evaluation Across Heterogeneous Tasks
The method is evaluated on three distinct settings:
-
A controlled synthetic Gaussian-mixture task, which allows inspection of
interpolation and extrapolation behavior.
DAR achieves thehighest Spearman correlation among the LLM-based methods,
following the target structure more closely, especially in extrapolation regions. -
Code performance prediction (KBSS latency and APPS memory). DAR obtains the
strongest overall performance,
showing gains in Spearman correlation (up to a 6-point improvement on KBSS) while maintaining competitive or lower MAE compared to baselines like SFT and pointwise RL rewards. -
Molecular property prediction on MoleculeNet using only SMILES strings. DAR consistently outperforms other LLM training strategies, achieving
statistically significant gains
on FreeSolv and Lipophilicity, suggesting thatdistribution-level optimization provides more reliable regression behavior than pointwise objectives alone.
Distributional Analysis and Diagnostics
The paper analyzes the resulting predictive distributions to assess their quality:
-
Predictive uncertainty is evaluated using the mean of sampled predictions as the point estimate and the standard deviation as the score. DAR shows
the strongest log-scale correlation between predictive standard deviation and prediction error,
indicating that its rollout dispersionbetter reflects example-level prediction difficulty.
-
During training under GRPO, DAR is shown to encourage a more useful exploration: it achieves a
higher bracketing rate
(meaning rollouts more often cover the ground-truth value) and maintainshigher rollout entropy during training,
suggesting itmitigates diversity collapse.
Limitations and Broader Impact
The work acknowledges limitations, noting that LLM-based regression still requires careful prompting and decoding choices, which can introduce variability. However, the paper concludes that DAR demonstrates that optimizing predictive distributions improves ranking quality and supports uncertainty estimation,
suggesting it is a promising direction for LLM-based scientific regression
by providing a foundation for using LLMs as flexible regressors.
Task Prompts Used
The experiments utilize specific system and instruction prompts tailored to the task:
(Examples provided in Section H)
-
Synthetic Distributional Regression: System Prompt asks to
predict the target value y
given input x, with an output format requiring a boxed numeric value. -
Coding Performance Prediction (APPS/KBSS): System Prompts are tailored to predict
peak memory usage in bytes
orkernel execution latency in milliseconds,
respectively, using the code/problem context as the instruction prompt. -
MoleculeNet Property Prediction (ESOL/FreeSolv/Lipo): System Prompts ask to predict specific properties like
aqueous solubility (logS)
given a SMILES string input.
Comparison to Baselines
DAR is compared against several baselines, including:
Improvements for AI systems
Here are specific improvements that can be made to existing AI systems by implementing the concepts from this paper:
-
Improving Predictive Distribution Quality for Regression Tasks: The core improvement is moving from optimizing individual point estimates (like standard MSE) to optimizing the entire predictive distribution formed by multiple samples. This addresses the fundamental issue of poorly calibrated and overly narrow predictions common in current LLM regression.
-
Enhancing Candidate Ranking Reliability: By rewarding rollouts based on their marginal contribution to the overall predictive distribution quality (using the Continuous Ranked Probability Score, CRPS), the resulting models will produce better-calibrated uncertainty estimates and more accurate relative ordering of candidates.
-
Improving Extrapolation Performance: The method is shown to be particularly effective in extrapolation regions (as seen in the synthetic task results). This means LLMs can make more reliable predictions for inputs that fall outside the training distribution, which is critical for real-world scientific discovery where novel compounds or complex code structures are encountered.
-
Increasing Predictive Diversity During Training: The Distribution-Aware Reward (DAR) explicitly encourages
useful spread
and preventsrollout diversity collapse.
This means the model will maintain a richer set of predictions during reinforcement learning training, leading to more robust exploration and better coverage of the target value range, rather than collapsing to a single mean-seeking output. -
Enabling More Robust Uncertainty Estimation: The paper demonstrates that the resulting rollout dispersion is better aligned with actual prediction errors (strong log-Pearson correlation). This allows downstream decision-making systems to use the model's variance not just as a generic measure of uncertainty, but as a more informative and calibrated estimate of predictive difficulty.
-
Improving Code Performance Prediction: For tasks like predicting Triton kernel latency or Python peak memory usage, applying DAR is shown to yield the strongest rank correlation gains (up to 4-6 points on KBSS), suggesting that the model will be significantly better at ranking code implementations by their true performance rather than just predicting an average error.
-
Improving Molecular Property Prediction: For predicting properties like ESOL or FreeSolv from SMILES strings, DAR outperforms other training strategies, leading to superior CRPS and WIS scores, meaning the model's predictions will be both more accurate and better calibrated regarding their uncertainty.
These improvements result in AI systems that are not just good at guessing a single number (pointwise accuracy) but are capable of producing reliable probability distributions that accurately reflect the model's own confidence, making them far more trustworthy for scientific screening and decision support.
Sources
- Regression Language Models for Code
- Scaling Open-Ended Reasoning to Predict the Future
- Beyond Token-level Supervision: Unlocking the Potential of Decoding-based Regression via Reinforcement Learning
- The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models
- Mol-Instructions: A Large-Scale Biomolecular Instruction Dataset for Large Language Models
- Uncertainty Quantification for Regression using Proper Scoring Rules
- Reasoning on Time-Series for Financial Technical Analysis
- Time Series Forecasting via Reasoning: A Slow-Thinking Approach with Reinforcement Fine-Tuned LLMs
- Language Model Embeddings Can Be Sufficient for Bayesian Optimization
- Training language models to follow instructions with human feedback
- Anticipatory Evaluation of Language Models
- Entropy-Preserving Reinforcement Learning
- Reaching Beyond the Mode: RL for Distributional Reasoning in Language Models
- Qwen2 Technical Report
- Proximal Policy Optimization Algorithms
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- HybridFlow: A Flexible and Efficient RLHF Framework
- OmniPred: Language Models as Universal Regressors
- Understanding LLM Embeddings for Regression
- Reasoning-Intensive Regression
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks