Tail-Likelihood Reinforcement Learning
cs.LG, stat.ML
Submitted: 2026-09-02
Updated: 2026-09-09
Project page: https://zanette-labs.github.io/TailRL-website
Terminology
Sources
- Qwen2.5-VL Technical Report
- Kevin: Multi-Turn RL for Generating CUDA Kernels
- Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
- Evaluating Large Language Models Trained on Code
- Self-Evolving Curriculum for LLM Reasoning
- Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models
- Robust Reinforcement Learning with Distributional Risk-averse formulation
- Training Verifiers to Solve Math Word Problems
- The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models
- CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models
- What is the objective of reasoning with reinforcement learning?
- StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback
- Differential Smoothing Mitigates Sharpening and Improves LLM Reasoning
- KL-Regularized Reinforcement Learning is Designed to Mode Collapse
- Target Policy Optimization
- CodeRL: Mastering Code Generation through Pretrained Models and Deep Reinforcement Learning
- Let's Verify Step by Step
- RLTF: Reinforcement Learning from Unit Test Feedback
- The gem5 Simulator: Version 20.0+
- The Reasoning Boundary Paradox: How Reinforcement Learning Constrains Language Models
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks