Messier: A High-Resolution Corpus for Cross-Benchmark Agent Evaluation
cs.AI, cs.DB
Submitted: 2026-07-28
Updated: 2026-08-30
Terminology
Sources
- RewardHackingAgents: Benchmarking Evaluation Integrity for LLM ML-Engineering Agents
- $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment
- Finance Agent Benchmark: Benchmarking LLMs on Real-world Financial Research Tasks
- AstaBench: Rigorous Benchmarking of AI Agents with a Scientific Research Suite
- ACEBench: Who Wins the Match Point in Tool Usage?
- SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
- DABstep: Data Agent Benchmark for Multi-step Reasoning
- EvilGenie: A Reward Hacking Benchmark
- Agent psychometrics: Task-level performance prediction in agentic coding benchmarks
- A Rosetta Stone for AI Benchmarks
- Steering Evaluation-Aware Language Models to Act Like They Are Deployed
- Holistic Agent Leaderboard: The Missing Infrastructure for AI Agent Evaluation
- SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
- Evaluating Test-Time Scaling of General LLM Agents
- Holistic Evaluation of Language Models
- BRIDGE: Predicting Human Task Completion Time From Model Performance
- AgentRewardBench: Evaluating Automatic Evaluations of Web Agent Trajectories
- Quantifying Variance in Evaluation Benchmarks
- Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
- Large Language Models Often Know When They Are Being Evaluated
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection