Pinned and Still Unstable: Within-Judge Verdict Variance and the Noise Floor of LLM-as-Judge Leaderboards
cs.CL
Submitted: 2026-09-27
Updated: 2026-09-30
Terminology
Sources
- Non-Determinism of "Deterministic" LLM Settings
- Position: Don't Use the CLT in LLM Evals With Fewer Than a Few Hundred Datapoints
- Humans or LLMs as the Judge? A Study on Judgement Biases
- Safer or Luckier? LLMs as Safety Evaluators Are Not Robust to Artifacts
- How is ChatGPT's behavior changing over time?
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback
- Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
- LZ Penalty: An information-theoretic repetition penalty for autoregressive language models
- Do Repetitions Matter? Strengthening Reliability in LLM Evaluations
- From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline
- Aligning with Human Judgement: The Role of Pairwise Preference in Large Language Model Evaluators
- LLM Evaluators Recognize and Favor Their Own Generations
- Verbosity Bias in Preference Labeling by Large Language Models
- Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge
- Large Language Models are Inconsistent and Biased Evaluators
- Let Me Speak Freely? A Study on the Impact of Format Restrictions on Performance of Large Language Models
- Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges
- The Verification Tax: Fundamental Limits of AI Auditing in the Rare-Error Regime
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering