Tandem Reinforcement Learning with Verifiable Rewards
cs.AI
Submitted: 2026-06-26
Updated: 2026-09-27
Code: https://github.com/CSSLab/Tandem-RLVR
Terminology
Sources
- Constitutional AI: Harmlessness from AI Feedback
- Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision
- Evaluating Large Language Models Trained on Code
- GPG: A Simple and Strong Reinforcement Learning Baseline for Model Reasoning
- SSR: Speculative Parallel Scaling Reasoning in Test-time
- The Collaboration Gap: Exploration and Benchmarking of Open-World Agentic Cooperation
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Designing Skill-Compatible AI: Methodologies and Frameworks in Chess
- Measuring Mathematical Problem Solving With the MATH Dataset
- Distilling the Knowledge in a Neural Network
- The Steganographic Potentials of Language Models
- Prover-Verifier Games improve legibility of LLM outputs
- Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe
- Understanding R1-Zero-Like Training: A Critical Perspective
- Sparse but Critical: A Token-Level Analysis of Distributional Shifts in RLVR Fine-Tuning of LLMs
- Training language models to follow instructions with human feedback
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Large language models can learn and generalize steganographic chain-of-thought under process supervision
- Solving math word problems with process- and outcome-based feedback
- MARS: toward more efficient multi-agent collaboration for LLM reasoning
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection