When Is an Agent Evaluation Over? Outcome Finality and Cross-Unit Separation
cs.AI, cs.CY
Submitted: 2026-08-14
Updated: 2026-08-27
Terminology
Sources
- Declare and Justify: Explicit assumptions in AI evaluations are necessary for effective regulation
- Done, But Not Sure: Disentangling World Completion from Self-Termination in Embodied Agents
- Cordon: Semantic Transactions for Tool-Using LLM Agents
- AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents
- How Benchmarks Mis-Score Computer-Use Agents
- WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?
- Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation
- Scaling Agents for Computer Use
- Measurement and Fairness
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
- Training Agents to Self-Report Misbehavior
- Engagement Process: Rethinking the Temporal Interface of Action and Observation
- AgentCanary: A Security Evaluation Framework for Autonomous AI Agents in Real Executable Environments
- ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities
- What Twelve LLM Agent Benchmark Papers Disclose About Themselves: A Pilot Audit and an Open Scoring Schema
- Do Agent Benchmarks Measure Capability? Protocol Validity in the Age of Agentic AI
- AsyncTool: Evaluating the Asynchronous Function Calling Capability under Multi-Task Scenarios
- Audit Cards: Contextualizing AI Evaluations
- Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack
- Automated Benchmark Auditing for AI Agents and Large Language Models
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection