MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation
cs.CL, cs.AI
Submitted: 2025-10-21
Updated: 2026-08-27
Code: https://github.com/choics2623/MENTOR-RL
Project page: https://artofproblemsolving.com/wiki/index.php/2023_AMC_12A_Problems
Terminology
Sources
- ReSearch: Learning to Reason with Search for LLMs via Reinforcement Learning
- AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning
- Beyond Imitation: Learning Key Reasoning Steps from Dual Chain-of-Thoughts in Reasoning Distillation
- ReTool: Reinforcement Learning for Strategic Tool Use in LLMs
- Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models
- TPTU-v2: Boosting Task Planning and Tool Usage of Large Language Model-based Agents in Real-world Systems
- Solving Quantitative Reasoning Problems with Language Models
- LLMs Can Easily Learn to Reason from Demonstrations Structure, not content, is what matters!
- LoRA: Low-Rank Adaptation of Large Language Models
- R-Zero: Self-Evolving Reasoning LLM from Zero Data
- GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
- Self-Training Large Language Models for Tool-Use Without Demonstrations
- Student-Centered Distillation Narrows the Agentic Gap Between Small and Large LLMs
- ART: Automatic multi-step reasoning and tool-use for large language models
- Distilling LLM Agent into Small Models with Retrieval and Code Tools
- Climbing the Ladder of Reasoning: What LLMs Can-and Still Can't-Solve after SFT?
- Text Embeddings by Weakly-Supervised Contrastive Pre-training
- AgentDistill: Training-Free Agent Distillation with Generalizable MCP Boxes
- Qwen3 Technical Report
- OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering