Confident Rankings with Fewer Items: Adaptive LLM Evaluation with Continuous Scores
cs.CL, cs.AI
Submitted: 2026-01-20
Updated: 2026-09-14
Comments: EMNLP 2026 Main Conference
Code: https://github.com/trismik/continuous-cat
License: http://creativecommons.org/licenses/by/4.0/
The gist: Computerized Adaptive Testing (CAT) has proven effective for efficient LLM evaluation on multiple-choice benchmarks, but modern LLM evaluation increasingly relies on generation tasks where outputs
Terminology
Abstract
Computerized Adaptive Testing (CAT) has proven effective for efficient LLM evaluation on multiple-choice benchmarks, but modern LLM evaluation increasingly relies on generation tasks where outputs are scored continuously rather than marked correct/incorrect. We present a principled extension of IRT-based adaptive testing to continuous bounded scores (ROUGE, BLEU, LLM-as-a-Judge) by replacing the Bernoulli response distribution with a heteroskedastic normal distribution. Building on this, we introduce an uncertainty aware ranker with adaptive stopping criteria that achieves reliable model ranking while testing as few items and as cheaply as possible. We validate our method on five benchmarks spanning n-gram-based, embedding-based, and LLM-as-judge metrics. Our method improves ranking correlation by 0.13 τ over random sampling and has 99% accuracy on confident predictions while using 2% of the items after a one-time calibration step.
Sources
- Survey of Computerized Adaptive Testing: A Machine Learning Perspective
- G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment
- No Language Left Behind: Scaling Human-Centered Machine Translation
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering