Aggregate Disambiguation Systems
stat.ME, cs.AI
Submitted: 2026-08-31
Updated: 2026-08-31
Terminology
Sources
- RoPoLL: Robust Panel of LLM Judges
- SCOPE: Selective Conformal Optimized Pairwise LLM Judging
- Nine Judges, Two Effective Votes: Correlated Errors Undermine LLM Evaluation Panels
- Calibrate, Don't Curate: Label-Efficient Estimation from Noisy LLM Judges
- Who can we trust? LLM-as-a-jury for Comparative Assessment
- Judging the Judges: A Systematic Study of Position Bias in LLM-as-a-Judge
- Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models
- A Finite-Calibration Regime Map for LLM Judge Panels
Related papers
- Doubly robust inference via calibration
- Bayesian Empirical Bayes: Simultaneous Inference from Probabilistic Symmetries
- Flexible Nonparametric Inference for Causal Effects under the Front-Door Model
- Deployment of AI-Assisted Interventions: Capacity Constraints and Noisy Compliance
- A Survey on Archetypal Analysis
- Dynamic Spatial Bayesian Machine Learning Model: Applications to Intergenerational Economic Mobility and Geographic Income Inequality in the United States