What Makes a Terminal-Bench Task Hard? Separating Genuine Hardness from Fake-Hardness on an Adjudicated Agentic Corpus
cs.LG, cs.AI
Submitted: 2026-09-20
Updated: 2026-09-20
Terminology
Sources
- Agent psychometrics: Task-level performance prediction in agentic coding benchmarks
- Position: Coding Benchmarks Are Misaligned with Agentic Software Engineering
- LLMs Gaming Verifiers: RLVR can Lead to Reward Hacking
- Measuring Five-Nines Reliability: Sample-Efficient LLM Evaluation in Saturated Benchmarks
- On the limits and opportunities of AI reviewers: Reviewing the reviews of Nature-family papers with 45 expert scientists
- LLM-as-a-Verifier: A General-Purpose Verification Framework
- Auditing LLM Benchmarks with Item Response Theory
- Does SWE-Bench-Verified Test Agent Ability or Model Memory?
- Cross-Context Verification: Hierarchical Detection of Benchmark Contamination through Session-Isolated Analysis
- Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack
- Harness-Bench: Measuring Harness Effects across Models in Realistic Agent Workflows
- Hardening Agent Benchmarks with Adversarial Hacker-Fixer Loops
- Lost in Benchmarks? Rethinking Large Language Model Benchmarking with Item Response Theory
- Establishing Best Practices for Building Rigorous Agentic Benchmarks
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks