Mediocrity is the key for LLM as a Judge Anchor Selection
cs.CL
Submitted: 2026-03-17
Updated: 2026-09-02
Code: https://github.com/IBM/Anchor-Selection
Terminology
Sources
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- The Falcon Series of Open Language Models
- Mixtral of Experts
- Naturally Occurring Feedback is Common, Extractable and Useful
- Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
- From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline
- A Survey on LLM-as-a-Judge
- gpt-oss-120b & gpt-oss-20b Model Card
- Beyond Pairwise: Empowering LLM Alignment With Ranked Choice Modeling
- Gemma 3 Technical Report
- Zero-shot Benchmarking: A Framework for Flexible and Scalable Automatic Evaluation of Language Models
- TrustJudge: Inconsistencies of LLM-as-a-Judge and How to Alleviate Them
- Verbosity Bias in Preference Labeling by Large Language Models
- Qwen3 Technical Report
- Qwen2 Technical Report
- Qwen2.5 Technical Report
- Yi: Open Foundation Models by 01.AI
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering