HiDiffTIR: Hierarchical Difficulty-Aware Policy Optimization for Multi-Turn Tool-Integrated Reasoning
cs.CL, cs.AI
Submitted: 2026-08-22
Updated: 2026-09-01
Comments: Accepted by EMNLP 2026 (Findings)
Code: https://github.com/YucanGuo/HiDiffTIR
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning
- Asymmetric Prompt Weighting for Reinforcement Learning with Verifiable Rewards
- Understanding the planning of LLM agents: A survey
- Reinforcement Learning with Rubric Anchors
- ToRL: Scaling Tool-Integrated RL
- Understanding Tool-Integrated Reasoning
- Large Language Model Agent: A Survey on Methodology, Applications and Challenges
- The Landscape of Emerging AI Agent Architectures for Reasoning, Planning, and Tool Calling: A Survey
- GPT-4 Technical Report
- ToolRL: Reward is All Tool Learning Needs
- MatchTIR: Fine-Grained Supervision for Tool-Integrated Reasoning via Bipartite Matching
- SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning
- Qwen3 Technical Report
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- ToolAlpaca: Generalized Tool Learning for Language Models with 3000 Simulated Cases
- Solving math word problems with process- and outcome-based feedback
- VisualPRM: An Effective Process Reward Model for Multimodal Reasoning
- Feedback-Driven Tool-Use Improvements in Large Language Models via Automated Build Environments
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering