ARISE-RL: Agentic Rubric-Grounded Iterative Self-Evolution with Reinforcement Learning
cs.AI
Submitted: 2026-09-01
Updated: 2026-09-01
Code: https://github.com/Alibaba-NLP/qqr
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning
- DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents
- A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence
- VitaBench: Benchmarking LLM Agents with Versatile Interactive Tasks in Real-world Applications
- Distilling the Knowledge in a Neural Network
- R-Zero: Self-Evolving Reasoning LLM from Zero Data
- Agent0-VL: Exploring Self-Evolving Agent for Tool-Integrated Vision-Language Reasoning
- Search Self-play: Pushing the Frontier of Agent Capability without Supervision
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- ResearchRubrics: A Benchmark of Prompts and Rubrics For Evaluating Deep Research Agents
- Towards Agentic Self-Learning LLMs in Search Environment
- RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
- EvolveR: Self-Evolving LLM Agents through an Experience-Driven Lifecycle
- Qwen3 Technical Report
- ReAct: Synergizing Reasoning and Acting in Language Models
- On-Policy Context Distillation for Language Models
- Dr. Zero: Self-Evolving Search Agents without Training Data
- ArenaRL: Scaling RL for Open-Ended Agents via Tournament-based Relative Ranking
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection