TRACE: Diagnosing Verifier Brittleness in Agentic Evaluation
cs.CL
Submitted: 2026-10-08
Updated: 2026-10-08
Code: https://github.com/RGaonkar/trace-verifier-stress-tests
Terminology
Sources
- Constitutional AI: Harmlessness from AI Feedback
- $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment
- Measuring what Matters: Construct Validity in Large Language Model Benchmarks
- Training Verifiers to Solve Math Word Problems
- Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards
- Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking
- Beyond Math and Code: Lightweight Corpus-Grounded Process Rewards for Factual Question Answering
- EvilGenie: A Reward Hacking Benchmark
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
- AI Agents That Matter
- Catastrophic Goodhart: regularizing RLHF with KL divergence does not mitigate heavy-tailed reward misspecification
- AgentBench: Evaluating LLMs as Agents
- AgentRewardBench: Evaluating Automatic Evaluations of Web Agent Trajectories
- Adversarial NLI: A New Benchmark for Natural Language Understanding
- Training language models to follow instructions with human feedback
- Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
- Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning
- Learning to summarize from human feedback
- Reward Hacking Benchmark: Measuring Exploits in LLM Agents with Tool Use
- HiL-Bench (Human-in-Loop Benchmark): Do Agents Know When to Ask for Help?
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering