Reinforcement Learning over Predictive Distributions for LLM Regression

arXiv:2605.20740 · cs.LG, cs.AI, cs.CL · Submitted 2026-05-20 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Reinforcement Learning over Predictive Distributions for LLM Regression".

Jane: Large language models can predict real-valued quantities from heterogeneous inputs such as text, code, and molecular strings, but most training objectives score each decoded floating-point number independently,

Tom: First, who's behind it and why it matters.

Title and authors: Tom: Alright, so we're starting with the title and who's behind this paper, "Reinforcement Learning over Predictive Distributions for LLM Regression." It tells us immediately that the core idea is using reinforcement learning specifically to shape how these models create their probability distributions for regression tasks.

Jane: And the authors are Jungsoo Park, Hyungjoo Chae, Ethan Mendes, Jay DeYoung, Varsha Kishore, and Wei Xu. They’re a solid group of researchers from places like Georgia Institute of Technology and Allen Institute for AI.

Lu: I find it interesting that they focused on regression tasks specifically because predicting real-valued quantities is where we really need calibrated uncertainty estimates to make sense of the results.

Meng: When you look at the abstract, they point out that most current training objectives just score each decoded floating-point number independently, which means they improve those point estimates but don't guarantee a calibrated predictive distribution, which is a big problem for applications requiring ranking or uncertainty estimation.

Lalam: It seems like this paper is directly addressing that limitation by introducing Distribution-Aware Reward to train the language models to produce better distributions instead of just optimizing individual outputs against scalar targets.

The paper's summary: Tom: So, looking at what they actually propose in the paper, Jungsoo Park et al. introduce Distribution-Aware Reward as an on-policy reinforcement learning objective designed to train language models to produce better predictive distributions for regression tasks rather than just optimizing individual decoded outputs against scalar targets.

Jane: In simpler terms, they are treating all the different predictions the model makes during training as a collection, or an empirical predictive distribution, and they evaluate that whole collection using something called the Continuous Ranked Probability Score, which is a proper way to score distributions.

Lu: That CRPS evaluation is key because it gives them a principled way to check how good their resulting distributions are at actually predicting the target values.

Meng: And they assign credit for each individual rollout using something called leave-one-out credit assignment, which means they figure out how much each specific prediction contributes to making the whole distribution better, rewarding predictions that improve the dispersion around the ground truth.

Lalam: It’s about moving from optimizing single guesses to optimizing the quality of the entire set of guesses, which is a significant shift in what we train these models for.

The paper's improvements: Tom: One of the main improvements they highlight is that pointwise rewards, like standard MSE, score each rollout independently and actually encourage predictions to collapse toward the mean, which can lead to predictions that are poorly calibrated or overly narrow.

Jane: But Distribution-Aware Reward does something different; it scores each rollout by its leave-one-out contribution to the full predictive distribution, which rewards both accuracy and useful spread around the ground truth.

Lu: They show that this results in rollouts that are both accurate and better calibrated as a set, which improves both pointwise error and rank correlation, especially when looking at how they perform on interpolation and extrapolation behavior in synthetic Gaussian-mixture tasks.

Meng: This is interesting because it suggests that optimizing the predictive distribution formed by multiple LLM rollouts leads to more reliable regression behavior than just optimizing individual decoded outputs against scalar targets.

Lalam: It means we’re getting predictions that are not just a better single number, but a set of numbers whose collective behavior is much more trustworthy for real-world use.

Conclusion: Tom: So to wrap up on "Reinforcement Learning over Predictive Distributions for LLM Regression," the authors demonstrate that DAR optimizes the predictive distribution formed by multiple LLM rollouts, moving past optimizing individual decoded outputs to match scalar targets.

Jane: They conclude that this method encourages predictions that are both accurate and better calibrated as a set, which improves both pointwise error and rank correlation across different settings like synthetic tasks and molecular property prediction.

Lu: The analysis of the resulting predictive distributions also shows that the rollout dispersion better reflects example-level prediction difficulty, indicating their spread is more meaningful than just a generic measure of uncertainty.

Meng: They also observed that under training with GRPO, DAR encourages more useful exploration by achieving a higher bracketing rate and maintaining higher rollout entropy during training, which helps mitigate diversity collapse.

Lalam: Ultimately, the paper shows that optimizing these predictive distributions provides more reliable regression behavior than pointwise objectives alone, offering a promising direction for using LLMs as flexible regressors in scientific fields.

Georgia Institute of Technology · Allen Institute for AI

cs.LG, cs.AI, cs.CL

Submitted: 2026-05-20

Updated: 2026-10-06

Code: https://github.com/verl-project/verl

Importance score: 90/100

The gist: Large language models can predict real-valued quantities from heterogeneous inputs such as text, code, and molecular strings, but most training objectives score each decoded floating-point number

Key concepts

Continuous Ranked Probability Score (CRPS)
CRPS is a proper scoring rule used to measure how good a predicted probability distribution is compared to the true target value. By maximizing the negative CRPS, DAR encourages the model's predictions to be not just accurate but also appropriately dispersed, ensuring it captures uncertainty well.
Leave-One-Out Credit Assignment
This technique assigns credit for a prediction by comparing the score of the full set of predictions against a set where one specific prediction is removed. This helps determine how much each individual rollout contributes to the overall quality of the predictive distribution.
Pointwise Rewards vs. DAR
Pointwise rewards, like Mean Squared Error (MSE), treat each output independently, often leading to narrow or poorly calibrated predictions. DAR contrasts this by optimizing the collective predictive distribution, encouraging models to be accurate and better calibrated as a set.
Predictive Distribution
Instead of focusing on a single decoded number, DAR treats all multiple LLM rollouts from the same input as an empirical predictive distribution. This allows the objective function to evaluate how well the model's predictions cover the target space, leading to better uncertainty estimates.

Terminology

Summary

Large language models can predict real-valued quantities from heterogeneous inputs such as text, code, and molecular strings, but most training objectives score each decoded floating-point number independently, improving point estimates without ensuring calibrated predictive distributions.

The gist

Distribution-Aware Reward (DAR) is an on-policy reinforcement learning objective that trains language models to produce better predictive distributions for regression tasks by treating multiple decoded samples as an empirical predictive distribution and evaluating it with the Continuous Ranked Probability Score (CRPS).

How it works

  1. The method treats multiple decoded samples as an empirical predictive distribution.

  2. It evaluates this distribution against the target using the Continuous Ranked Probability Score (CRPS), which is defined as a proper scoring rule for predictive distributions. The objective is to maximize the negated CRPS, denoted as Si = −CRPS(Fi, yi).

  3. To assign credit to individual rollouts, DAR uses leave-one-out credit assignment. For each rollout prediction, it calculates the marginal contribution by comparing the score of the full distribution with a leave-one-out empirical distribution obtained by removing that specific rollout: R DAR ik = Si − S(-k) i.

  4. This mechanism rewards rollouts that are both accurate and appropriately dispersed, encouraging predictions that are accurate and better calibrated as a set, rather than just optimizing individual outputs against scalar targets.

Key Contributions to Optimization

(Based on the comparison between pointwise rewards and DAR)

  1. Pointwise rewards (e.g., MSE) score each rollout independently, encouraging predictions to collapse toward the mean. This can lead to predictions that are poorly calibrated, overly narrow, or one-sided, which weakens uncertainty estimation and rank-based evaluation.

  2. DAR optimizes the predictive distribution formed by multiple LLM rollouts, moving beyond optimizing individual decoded outputs to match scalar targets.

  3. DAR encourages predictions that are accurate and better calibrated as a set (Figure 1), improving both pointwise error and rank correlation.

Evaluation Across Heterogeneous Tasks

The method is evaluated on three distinct settings:

  1. A controlled synthetic Gaussian-mixture task, which allows inspection of interpolation and extrapolation behavior. DAR achieves the highest Spearman correlation among the LLM-based methods, following the target structure more closely, especially in extrapolation regions.

  2. Code performance prediction (KBSS latency and APPS memory). DAR obtains the strongest overall performance, showing gains in Spearman correlation (up to a 6-point improvement on KBSS) while maintaining competitive or lower MAE compared to baselines like SFT and pointwise RL rewards.

  3. Molecular property prediction on MoleculeNet using only SMILES strings. DAR consistently outperforms other LLM training strategies, achieving statistically significant gains on FreeSolv and Lipophilicity, suggesting that distribution-level optimization provides more reliable regression behavior than pointwise objectives alone.

Distributional Analysis and Diagnostics

The paper analyzes the resulting predictive distributions to assess their quality:

  1. Predictive uncertainty is evaluated using the mean of sampled predictions as the point estimate and the standard deviation as the score. DAR shows the strongest log-scale correlation between predictive standard deviation and prediction error, indicating that its rollout dispersion better reflects example-level prediction difficulty.

  2. During training under GRPO, DAR is shown to encourage a more useful exploration: it achieves a higher bracketing rate (meaning rollouts more often cover the ground-truth value) and maintains higher rollout entropy during training, suggesting it mitigates diversity collapse.

Limitations and Broader Impact

The work acknowledges limitations, noting that LLM-based regression still requires careful prompting and decoding choices, which can introduce variability. However, the paper concludes that DAR demonstrates that optimizing predictive distributions improves ranking quality and supports uncertainty estimation, suggesting it is a promising direction for LLM-based scientific regression by providing a foundation for using LLMs as flexible regressors.

Task Prompts Used

The experiments utilize specific system and instruction prompts tailored to the task:

(Examples provided in Section H)

  1. Synthetic Distributional Regression: System Prompt asks to predict the target value y given input x, with an output format requiring a boxed numeric value.

  2. Coding Performance Prediction (APPS/KBSS): System Prompts are tailored to predict peak memory usage in bytes or kernel execution latency in milliseconds, respectively, using the code/problem context as the instruction prompt.

  3. MoleculeNet Property Prediction (ESOL/FreeSolv/Lipo): System Prompts ask to predict specific properties like aqueous solubility (logS) given a SMILES string input.

Comparison to Baselines

DAR is compared against several baselines, including:

Improvements for AI systems

Here are specific improvements that can be made to existing AI systems by implementing the concepts from this paper:

  1. Improving Predictive Distribution Quality for Regression Tasks: The core improvement is moving from optimizing individual point estimates (like standard MSE) to optimizing the entire predictive distribution formed by multiple samples. This addresses the fundamental issue of poorly calibrated and overly narrow predictions common in current LLM regression.

  2. Enhancing Candidate Ranking Reliability: By rewarding rollouts based on their marginal contribution to the overall predictive distribution quality (using the Continuous Ranked Probability Score, CRPS), the resulting models will produce better-calibrated uncertainty estimates and more accurate relative ordering of candidates.

  3. Improving Extrapolation Performance: The method is shown to be particularly effective in extrapolation regions (as seen in the synthetic task results). This means LLMs can make more reliable predictions for inputs that fall outside the training distribution, which is critical for real-world scientific discovery where novel compounds or complex code structures are encountered.

  4. Increasing Predictive Diversity During Training: The Distribution-Aware Reward (DAR) explicitly encourages useful spread and prevents rollout diversity collapse. This means the model will maintain a richer set of predictions during reinforcement learning training, leading to more robust exploration and better coverage of the target value range, rather than collapsing to a single mean-seeking output.

  5. Enabling More Robust Uncertainty Estimation: The paper demonstrates that the resulting rollout dispersion is better aligned with actual prediction errors (strong log-Pearson correlation). This allows downstream decision-making systems to use the model's variance not just as a generic measure of uncertainty, but as a more informative and calibrated estimate of predictive difficulty.

  6. Improving Code Performance Prediction: For tasks like predicting Triton kernel latency or Python peak memory usage, applying DAR is shown to yield the strongest rank correlation gains (up to 4-6 points on KBSS), suggesting that the model will be significantly better at ranking code implementations by their true performance rather than just predicting an average error.

  7. Improving Molecular Property Prediction: For predicting properties like ESOL or FreeSolv from SMILES strings, DAR outperforms other training strategies, leading to superior CRPS and WIS scores, meaning the model's predictions will be both more accurate and better calibrated regarding their uncertainty.

These improvements result in AI systems that are not just good at guessing a single number (pointwise accuracy) but are capable of producing reliable probability distributions that accurately reflect the model's own confidence, making them far more trustworthy for scientific screening and decision support.

Sources

Related papers