How Many Humans Are 32 LLM Judges Worth?

arXiv:2609.21277 · cs.CL, cs.LG · Submitted 2026-09-18 · Read on arXiv

cs.CL, cs.LG

Submitted: 2026-09-18

Updated: 2026-09-24

Comments: 18 pages, 10 figures, and 8 tables. Code and data: https://github.com/Chao1208/chaosnli-judge-votes

Code: https://github.com/Chao1208/chaosnli-judge-votes

License: http://creativecommons.org/licenses/by/4.0/

The gist: How many human judgments does a panel of language models represent? The answer depends on what is matched.

Terminology

Abstract

How many human judgments does a panel of language models represent? The answer depends on what is matched. We audit categorical judge panels against empirical human label distributions, retaining disagreement that binary errors relative to one gold label collapse. We measure spectral residual diversity by matching the participation ratio of a normalized residual Gram matrix to conditionally independent human-reference draws, giving nu H. We separately match distributional squared error, giving nu MSE. Across three ChaosNLI tasks, the same 32-judge panels have nu H=4.24--6.50 but nu MSE=2.30--3.75. A spectral identity separates the eigenvalues, member energies, and averaging-direction weights that determine error. Realizable hard-label panels show that greater spectral diversity can accompany worse distribution recovery even with equal member energies and nonnegative correlations. In the observed panels, within-size ranking agreement varies sharply by task; some member additions produce conflicting changes that persist across two item halves. The consensus-direction share of centered residual variance is gamma co=43.8% on MNLI-m and 33.7% on SNLI, quantifying shared variation retained by averaging. We provide aligned votes and analysis protocols for auditing these distinctions. Effective size is therefore a target-specific measurement: spectral diversity and distribution recovery should not be treated as interchangeable measures of panel quality or as general human-replacement rates.

Related papers