ContinualSkillBench: Can LLM Agents Truly Evolve Their Capabilities?
Tianyi Guan, Yiding Wang, Haotong Yang, Siyuan Cao, Shirui Liu, Yi Hu, Jiaqi Li, Muhan Zhang
cs.AI, cs.CL, cs.LG
Submitted: 2026-08-04
Code: https://github.com/gtynnn060110-hash/continual-skill-bench-final
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- LEXam: Benchmarking Legal Reasoning on 340 Law Exams
- LawBench: Benchmarking Legal Knowledge of Large Language Models
- Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models
- LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models
- HealthBench: Evaluating Large Language Models Towards Improved Human Health
- Putnam-AXIOM: A Functional and Static Benchmark for Measuring Higher Level Mathematical Reasoning in LLMs
- SkillCraft: Can LLM Agents Learn to Use Tools Skillfully?
- CL-bench: A Benchmark for Context Learning
- OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems
- INVESTORBENCH: A Benchmark for Financial Decision-Making Tasks with LLM-based Agent
- OfficeQA Pro: An Enterprise Benchmark for End-to-End Grounded Reasoning
- Generalization to New Sequential Decision Making Tasks with In-Context Learning
- Toolformer: Language Models Can Teach Themselves to Use Tools
- SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
- HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face
- MedicalAgentsBench for Complex Medical Reasoning: Comparing Internalized Reasoning Models versus Externalized Agent-based Frameworks
- PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments
- Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
- Continual Learning for Large Language Models: A Survey
- WritingBench: A Comprehensive Benchmark for Generative Writing
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection