Ask Which, Not How Good: Sizing Benchmarks Scored by an LLM
cs.AI
Submitted: 2026-08-17
Updated: 2026-08-17
Terminology
Sources
- Diagnosing the Reliability of LLM-as-a-Judge via Item Response Theory
- When Judgment Becomes Noise: How Design Failures in LLM Judge Benchmarks Silently Undermine Validity
- Hidden Measurement Error in LLM Pipelines Distorts Annotation, Evaluation, and Benchmarking
- Reliability without Validity: A Systematic, Large-Scale Evaluation of LLM-as-a-Judge Models Across Agreement, Consistency, and Bias
- Exploring LLM Autoscoring Reliability in Large-Scale Writing Assessments Using Generalizability Theory
- The Coin Flip Judge? Reliability and Bias in LLM-as-a-Judge Evaluation
- When the Judge Changes, So Does the Measurement: Auditing LLM-as-Judge Reliability
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection