CARE: Confounder-Aware Aggregation for Reliable LLM Evaluation
cs.LG, cs.AI, stat.ML
Submitted: 2026-02-09
Updated: 2026-09-08
Comments: ICLM 2026
Code: https://github.com/SprocketLab/CARE
License: http://creativecommons.org/licenses/by/4.0/
The gist: LLM-as-a-judge ensembles are the standard paradigm for scalable evaluation, but their aggregation mechanisms suffer from a fundamental flaw: they implicitly assume that judges provide independent
Terminology
Abstract
LLM-as-a-judge ensembles are the standard paradigm for scalable evaluation, but their aggregation mechanisms suffer from a fundamental flaw: they implicitly assume that judges provide independent estimates of true quality. However, in practice, LLM judges exhibit correlated errors caused by shared latent confounders -- such as verbosity, stylistic preferences, or training artifacts -- causing standard aggregation rules like majority vote or averaging to provide little gain or even amplify systematic mistakes. To address this, we introduce CARE, a confounder-aware aggregation framework that explicitly models LLM judge scores as arising from both a latent true-quality signal and shared confounding factors. Rather than heuristically re-weighting judges, CARE separates quality from confounders without access to ground-truth labels. We provide theoretical guarantees for identifiability and finite-sample recovery under shared confounders, and we quantify the systematic bias incurred when aggregation models omit confounding latent factors. Across 12 public benchmarks spanning continuous scoring, binary classification, and pairwise preference settings, CARE improves aggregation accuracy, reducing error by up to 26.8%. Code is released in https://github.com/SprocketLab/CARE.
Sources
- Yi: Open Foundation Models by 01.AI
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
- Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation
- GPT-4o System Card
- PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
- Gemma 3 Technical Report
- Debating with More Persuasive LLMs Leads to More Truthful Answers
- Gemma 2: Improving Open Language Models at a Practical Size
- RewardBench: Evaluating Reward Models for Language Modeling
- Preference Leakage: A Contamination Problem in LLM-as-a-judge
- Towards Acyclic Preference Evaluation of Language Models via Multiple Evaluators
- CalibraEval: Calibrating Prediction Distribution to Mitigate Selection Bias in LLMs-as-Judges
- LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods
- The Llama 3 Herd of Models
- Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs
- JudgeBlender: Ensembling Judgments for Automatic Relevance Assessment
- Judging the Judges: A Systematic Study of Position Bias in LLM-as-a-Judge
- Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
- Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models
- Improving LLM-as-a-Judge Inference with the Judgment Distribution
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks