DE-Venus: A Data-Efficient RLVR Framework for Large Language Models
cs.LG
Submitted: 2026-09-03
Updated: 2026-09-03
Code: https://github.com/ZJU-DIG/DE-Venus
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- The Unreasonable Effectiveness of Entropy Minimization in LLM Reasoning
- CoT-Kinetics: A Theoretical Modeling Assessing LRM Reasoning Process
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
- Measuring Mathematical Problem Solving With the MATH Dataset
- REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization
- Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model
- Look Before You Leap: An Exploratory Study of Uncertainty Measurement for Large Language Models
- OpenAI o1 System Card
- Scalable Best-of-N Selection for Large Language Models via Self-Certainty
- LIMR: Less is More for RL Scaling
- Understanding R1-Zero-Like Training: A Critical Perspective
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Kimi k1.5: Scaling Reinforcement Learning with LLMs
- Latent Space Chain-of-Embedding Enables Output-free LLM Self-Evaluation
- VerifyBench: Benchmarking Reference-based Reward Systems for Large Language Models
- Qwen3 Technical Report
- TraPO: A Semi-Supervised Reinforcement Learning Framework for Boosting LLM Reasoning
- Can LLMs Learn to Reason Robustly under Noisy Supervision?
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks