Critical-State RL: Diagnosing Trainable States for Multi-Turn Tool Use
cs.LG, cs.CL
Submitted: 2026-09-21
Updated: 2026-09-21
Comments: 31 pages, 8 figures, 7 tables
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- PivotRL: High Accuracy Agentic Post-Training at Low Compute Cost
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- DAPO: An Open-Source LLM Reinforcement Learning System at Scale
- Reinforcing Multi-Turn Reasoning in LLM Agents via Turn-Level Reward Design
- TCPO: Turn-Level Credit Policy Optimization
- Credit Without Ground Truth: Auditing Step-Level Credit Assignment in LLM Agents Against Executed Replay
- Agent Lightning: Train ANY AI Agents with Reinforcement Learning
- ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL
- AgentPRM: Process Reward Models for LLM Agents via Step-Wise Promise and Progress
- SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks
- TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents
- CAST: Game Solvers as Turn-Level Teachers for LLM Agents
- LOTAPO: Leave-One-Turn Attribution for Self-Generated Process Rewards in Multi-Turn Search Reasoning
- Teach the Magnitude, Not the Direction: Verifier-Bounded Credit Assignment for Multi-Turn Multi-step LLM Agents
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- Solving math word problems with process- and outcome-based feedback
- Gemma 4 Technical Report
- Qwen3 Technical Report
- APIGen: Automated Pipeline for Generating Verifiable and Diverse Function-Calling Datasets
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks