Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms
cs.AI, cs.CL
Submitted: 2026-09-23
Updated: 2026-09-23
Terminology
Sources
- Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents
- AutoForge: Automated Environment Synthesis for Agentic Reinforcement Learning
- Web-Shepherd: Advancing PRMs for Reinforcing Web Agents
- Beyond Itinerary Planning-A Real-World Benchmark for Multi-Turn and Tool-Using Travel Tasks
- Training Verifiers to Solve Math Word Problems
- TextWorld: A Learning Environment for Text-based Games
- From Self-Evolving Synthetic Data to Verifiable-Reward RL: Post-Training Multi-turn Interactive Tool-Using Agents
- Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains
- MemoryArena: Benchmarking Agent Memory in Interdependent Multi-Session Agentic Tasks
- Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model
- OR-Gym: A Reinforcement Learning Library for Operations Research Problems
- Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
- Measuring AI Ability to Complete Long Software Tasks
- Let's Verify Step by Step
- LLM+P: Empowering Large Language Models with Optimal Planning Proficiency
- Agentic Reinforcement Learning with Implicit Step Rewards
- Mock Worlds, Real Skills: Building Small Agentic Language Models with Synthetic Tasks, Simulated Environments, and Rubric-Based Rewards
- Reliability without Validity: A Systematic, Large-Scale Evaluation of LLM-as-a-Judge Models Across Agreement, Consistency, and Bias
- SlopCodeBench: Benchmarking How Coding Agents Degrade Over Long-Horizon Iterative Tasks
- Towards a Science of AI Agent Reliability
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection