Clean Engineering, Unstable Measurement: A Preregistered Reliability Failure of Black-Box LLM Observers on Shared Endpoints
cs.AI, cs.LG
Submitted: 2026-09-03
Updated: 2026-09-03
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Non-Determinism of "Deterministic" LLM Settings
- LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks
- Lessons from the Trenches on Reproducible Evaluation of Language Models
- How is ChatGPT's behavior changing over time?
- Test Before You Deploy: Governing Updates in the LLM Supply Chain
- Diagnosing the Reliability of LLM-as-a-Judge via Item Response Theory
- The Benchmarking Epistemology: Validity Theory for Evaluating Machine Learning Models
- A Survey on LLM-as-a-Judge
- Rating Roulette: Self-Inconsistency in LLM-As-A-Judge Frameworks
- Measurement and Fairness
- LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods
- Reproducibility in NLP: What Have We Learned from the Checklist?
- Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations
- The Great Pretender: A Stochasticity Problem in LLM Jailbreak
- Reliability without Validity: A Systematic, Large-Scale Evaluation of LLM-as-a-Judge Models Across Agreement, Consistency, and Bias
- An Empirical Study of the Non-determinism of ChatGPT in Code Generation
- LLM Evaluators Recognize and Favor Their Own Generations
- Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
- Judging the Judges: A Systematic Study of Position Bias in LLM-as-a-Judge
- A Two-Sided Discussion of Preregistration of NLP Research
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection