How to Run Statistics over LLM Judges and Trust the Results: Calibrated Inference for Small-Sample AI Evaluation with evalstats
cs.CL, cs.AI, cs.HC, stat.ME
Submitted: 2026-09-20
Updated: 2026-09-20
Code: https://github.com/ianarawjo/evalstats
Project page: https://statsforevals.com
Terminology
Sources
- PPI++: Efficient Prediction-Powered Inference
- Guidelines for Empirical Studies in Software Engineering involving Large Language Models
- Power Analysis for Prediction-Powered Inference
- Bias and Uncertainty in LLM-as-a-Judge Estimation
- Which Metrics Save the Most Human Annotation? Prediction-Powered Evaluation and Meta-Evaluation
- Do Repetitions Matter? Strengthening Reliability in LLM Evaluations
- This human study did not involve human subjects: Validating LLM simulations as behavioral evidence
- Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations
- Programming by Chat: A Large-Scale Behavioral Analysis of 11,579 Real-World AI-Assisted IDE Sessions
- Generalized Prediction-Powered Inference, with Application to Binary Classifier Evaluation
- A Note on the Prediction-Powered Bootstrap
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering