Certified Selective Automation of LLM Agent Evaluation
cs.CL
Submitted: 2026-09-28
Updated: 2026-09-28
Terminology
Sources
- SCOPE: Selective Conformal Optimized Pairwise LLM Judging
- Qwen3-VL Technical Report
- $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment
- AutoEval Done Right: Using Synthetic Data for Model Evaluation
- The BrowserGym Ecosystem for Web Agent Research
- WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?
- LoRA: Low-Rank Adaptation of Large Language Models
- Model-agnostic Selective Labeling with Provable Statistical Guarantees
- Language Models (Mostly) Know What They Know
- Reinforcement Learning on Web Interfaces Using Workflow-Guided Exploration
- AgentRewardBench: Evaluating Automatic Evaluations of Web Agent Trajectories
- Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
- Autonomous Evaluation and Refinement of Digital Agents
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Self-Taught Evaluators
- An Illusion of Progress? Assessing the Current State of Web Agents
- Qwen3 Technical Report
- Self-Rewarding Language Models
- Agent-as-a-Judge: Evaluate Agents with Agents
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering