Toward Robust LLM-Based Judges: Taxonomic Bias Evaluation and Debiasing Optimization
cs.CL
Submitted: 2026-03-09
Updated: 2026-09-29
Code: https://github.com/hiyouga/EasyR1
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Lost in Benchmarks? Rethinking Large Language Model Benchmarking with Item Response Theory
- A Comprehensive Survey of Reward Models: Taxonomy, Applications, Challenges, and Future
- RM-Distiller: Exploiting Generative LLM for Reward Model Distillation
- JudgeLRM: Large Reasoning Models as a Judge
- Skywork-Reward-V2: Scaling Preference Data Curation via Human-AI Synergy
- Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment
- Making Bias Non-Predictive: Training Robust LLM Reasoning via Reinforcement Learning
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- Are We on the Right Way to Assessing LLM-as-a-Judge?
- RewardBench 2: Advancing Reward Model Evaluation
- Kimi K2: Open Agentic Intelligence
- Qwen3 Technical Report
- Qwen2.5 Technical Report
- Atla Selene Mini: A General Purpose Evaluation Model
- Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs
- Think-J: Learning to Think for Generative LLM-as-a-Judge
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Representation Learning with Contrastive Predictive Coding
- The Llama 3 Herd of Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering