PAST-Bench: Benchmarking the Foundations of Recursive Self-Improvement in Personal Agents
Shuhan Xue, Zixin Ding, Yichen Shen, Yinjie Wang, Zhenfei Yin, Yingcheng Wu, Yuxin Chen, Mengdi Wang, Ling Yang
cs.CL
Submitted: 2026-08-04
Comments: Code: https://github.com/Gen-Verse/PAST-Bench
Code: https://github.com/Gen-Verse/PAST-Bench
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- AgentArch: A Comprehensive Benchmark to Evaluate Agent Architectures in Enterprise
- A Comprehensive Survey of Self-Evolving AI Agents: A New Paradigm Bridging Foundation Models and Lifelong Agentic Systems
- A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence
- From Raw Experience to Skill Consumption: A Systematic Study of Model-Generated Agent Skills
- Recursive Harness Self-Improvement
- SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
- ATBench: A Diverse and Realistic Agent Trajectory Benchmark for Safety Evaluation and Diagnosis
- Beyond Semantic Similarity: Rethinking Retrieval for Agentic Search via Direct Corpus Interaction
- MUSE-Autoskill: Self-Evolving Agents via Skill Creation, Memory, Management, and Evaluation
- SkillOS: Learning Skill Curation for Self-Evolving Agents
- Self-Improvements in Modern Agentic Systems: A Survey
- OpenClaw-RL: Train Any Agent Simply by Talking
- Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory
- SkillOpt: Executive Strategy for Self-Evolving Agent Skills
- MemRL: Self-Evolving Agents via Runtime Reinforcement Learning on Episodic Memory
- ClawBench: Can AI Agents Complete Everyday Online Tasks?
- LifelongAgentBench: Evaluating LLM Agents as Lifelong Learners
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering