AhaBench: Do Agents Turn Experience into Reusable Insights? A Long-Horizon Benchmark for Continual Learning
cs.LG, cs.CL
Submitted: 2026-06-30
Updated: 2026-09-21
Comments: 37 pages; In submission to TMLR
License: http://creativecommons.org/licenses/by/4.0/
The gist: Modern language agents are expected to operate over long horizons: they ask follow-up questions, reuse worked examples, handle tool feedback, and adapt to delayed consequences.
Terminology
Abstract
Modern language agents are expected to operate over long horizons: they ask follow-up questions, reuse worked examples, handle tool feedback, and adapt to delayed consequences. Most evaluations still reset the agent after a prompt or score only the final state of one trajectory. AhaBench asks a more operational question: when a fixed model receives useful experience, does its later behavior improve under a related evaluation condition where the obvious support has been removed, changed, or delayed? The suite contains three components. Aha-Puzzle tests no-hint exploration after solved hidden-state puzzles; Aha-Euler turns Project-Euler-style mathematical ideas into generated taught/held-out tasks with exact validators; and Aha-Vending, an open-source implementation inspired by Vending-Bench, tests whether a simulated vending agent remains profitable while handling delayed feedback and operational incidents. AhaBench reports a three-part scorecard: Initial Score measures starting competence, Post-Experience Score measures the later empirical outcome, and Learning Lift is their difference. This decomposition is the main empirical message: models that use visible support well, models that reach high post-experience scores, and models that improve most during a run are not always the same. On the common eight-model panel, Claude Opus 4.6 leads aggregate Post-Experience Score at 64.3 and aggregate Learning Lift at +25.8, with Gemini 3.1 Pro close behind at 63.4. The component results explain the split: puzzle traces raise supported scores but often fail to become no-hint exploration behavior; Aha-Euler full teaching reaches 78.6-100.0% while answer-only transfer ranges from 0.0 to 73.9%; and Aha-Vending separates profitable incident handling from bankruptcy and no-order failure. We release benchmark tasks, rubrics, validators, simulator code, and interfaces for evaluating new agents.
Sources
- Program Synthesis with Large Language Models
- Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents
- MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
- Evaluating Large Language Models Trained on Code
- VeRA: Renewing Reasoning Benchmarks with Executable Specifications
- On the Measure of Intelligence
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Training Verifiers to Solve Math Word Problems
- BRAINTEASER: Lateral Thinking Puzzles for Large Language Models
- Let's Verify Step by Step
- AgentBench: Evaluating LLMs as Agents
- GAIA: a benchmark for General AI Assistants
- WebGPT: Browser-assisted question-answering with human feedback
- Gorilla: Large Language Model Connected with Massive APIs
- Progressive Neural Networks
- Reflexion: Language Agents with Verbal Reinforcement Learning
- Voyager: An Open-Ended Embodied Agent with Large Language Models
- $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
- WebArena: A Realistic Web Environment for Building Autonomous Agents
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks