ACLArena: Agent Continue Learning in Multi-stage Post-training
cs.AI
Submitted: 2026-09-21
Updated: 2026-09-21
Comments: 25 pages, 7 figures, under review
Code: https://github.com/WillDreamer/ACLArena
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes
- $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment
- DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
- Beyond Benchmarks: MathArena as an Evaluation Platform for Mathematics with LLMs
- Are We Done with MMLU?
- MiniLLM: On-Policy Distillation of Large Language Models
- Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
- TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension
- Dense Passage Retrieval for Open-Domain Question Answering
- Measuring and Narrowing the Compositionality Gap in Language Models
- GPQA: A Graduate-Level Google-Proof Q&A Benchmark
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- FIRST: Teach A Reliable Large Language Model Through Efficient Trustworthy Distillation
- GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models
- Qwen3 Technical Report
- MuSiQue: Multihop Questions via Single-hop Question Composition
- Text Embeddings by Weakly-Supervised Contrastive Pre-training
- ARLArena: A Unified Framework for Stable Agentic Reinforcement Learning
- MiMo-V2-Flash Technical Report
- HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection