An Empirical Study of Automating Agent Evaluation
cs.CL
Submitted: 2026-05-12
Updated: 2026-09-23
Code: https://github.com/awslabs/Agent-EvalKit
Terminology
Sources
- ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate
- Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation
- Evaluating Large Language Models Trained on Code
- ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
- TRAJECT-Bench:A Trajectory-Aware Benchmark for Evaluating Agentic Tool Use
- Beyond the Final Answer: Evaluating the Reasoning Trajectories of Tool-Augmented Agents
- API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
- Let's Verify Step by Step
- AgentBench: Evaluating LLMs as Agents
- G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment
- ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs
- Identifying the Risks of LM Agents with an LM-Emulated Sandbox
- JudgeBench: A Benchmark for Evaluating LLM-based Judges
- A Survey on Large Language Model based Autonomous Agents
- OpenHands: An Open Platform for AI Software Developers as Generalist Agents
- DeepResearchEval: An Automated Framework for Deep Research Task Construction and Agentic Evaluation
- The Rise and Potential of Large Language Model Based Agents: A Survey
- TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks
- Survey on Evaluation of LLM-based Agents
- Deep Research: A Survey of Autonomous Research Agents
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering