Reasoning and Tool-use Compete in Agentic RL:From Quantifying Interference to Disentangled Tuning
cs.AI
Submitted: 2026-02-01
Updated: 2026-08-31
Terminology
Sources
- Constitutional AI: Harmlessness from AI Feedback
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- VerlTool: Towards Holistic Agentic Reinforcement Learning with Tool Use
- Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning
- ReTool: Reinforcement Learning for Strategic Tool Use in LLMs
- MixLoRA: Enhancing Large Language Models Fine-Tuning with LoRA-based Mixture of Experts
- When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories
- WebSailor-V2: Bridging the Chasm to Proprietary Agents via Synthetic Data and Scalable Reinforcement Learning
- Towards a Theoretical Understanding to the Generalization of RLHF
- MoELoRA: Contrastive Learning Guided Mixture of Experts on Parameter-Efficient Fine-Tuning for Large Language Models
- ToolRL: Reward is All Tool Learning Needs
- WebResearcher: Unleashing unbounded reasoning capability in Long-Horizon Agents
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Agentic Reasoning and Tool Integration for LLMs via Reinforcement Learning
- Qwen2 Technical Report
- Text Embeddings by Weakly-Supervised Contrastive Pre-training
- AutoTIR: Autonomous Tools Integrated Reasoning via Reinforcement Learning
- DeepAnalyze: Agentic Large Language Models for Autonomous Data Science
- Reward-SQL: Boosting Text-to-SQL via Stepwise Execution-Aware Reasoning and Process-Supervised Rewards
- MaskSearch: A Universal Pre-Training Framework to Enhance Agentic Search Capability
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection