The Statistical Benefits of Multiple Responses for Learning from Demonstrations
stat.ML, cs.IT, cs.LG, math.IT, math.ST, stat.TH
Submitted: 2026-09-27
Updated: 2026-09-27
Terminology
Sources
- The Best of N Worlds: Aligning Reinforcement Learning with Best-of-N Sampling via max@k Optimisation
- Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
- Evaluating Large Language Models Trained on Code
- Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models
- Learning with Multiple Correct Answers -- Regret Bounds under Different Feedback Models
Related papers
- Behavior of prediction performance metrics with rare events
- Optimal Estimation of Generic Dynamics by Path-Dependent Neural Jump ODEs
- A Posterior-Dynamics Framework for Imaging Inverse Problems with Pretrained Diffusion Priors
- One Permutation Is All You Need: Fast, Deterministic Feature Importance and Model Stress-Testing
- Online Conformal Prediction for Non-Exchangeable Panel Data
- Deep Time-Series Forecasting in 10 Years: A Survey