Cherry-pick Override: LLM Judges Under-use the Non-Directional Verdicts Their Contract Authorizes
cs.SE, cs.AI, cs.CL, cs.MA
Submitted: 2026-06-05
Updated: 2026-09-12
Comments: Substantially revised and extended. The failure covers both non-directional verdicts the contract authorizes, not conflict alone. New probes separate recognition from commitment and bound the rates from below. The planned two-reviewer audit was not completed; all rates stay dataset-defined. 4 judges, 2 substrates. Code: https://github.com/HrxuAlbert/cherry-pick-override
Code: https://github.com/HrxuAlbert/cherry-pick-override
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- A Gentle Introduction to Conformal Prediction and Distribution-Free Uncertainty Quantification
- Improving Factuality and Reasoning in Language Models through Multiagent Debate
- Not Wrong, But Untrue: LLM Overconfidence in Document-Based Queries
- No Certificate, No Categorical Speech Act: A Brouwerian Assertibility Constraint for Public Reason
- Trust or Escalate: LLM Judges with Provable Guarantees for Human Agreement
- Multi-Agent Verification: Scaling Test-Time Compute with Multiple Verifiers
- G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment
- AVeriTeC: A Dataset for Real-world Claim Verification with Evidence from the Web
- Get Your Vitamin C! Robust Fact Verification with Contrastive Evidence
- Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models
- From Debate to Decision: Conformal Social Choice for Safe Multi-Agent Deliberation
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Related papers
- Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
- GitSkills: A Dataset of Agent Skills on GitHub
- SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
- PackMonitor: Enabling Zero Package Hallucinations Through Decoding-Time Monitoring
- IntentCoding: Amplifying User Intent in Code Generation
- Incentives and Outcomes in Bug Bounties