Which Metrics Save the Most Human Annotation? Prediction-Powered Evaluation and Meta-Evaluation
cs.CL, stat.ML
Submitted: 2026-08-27
Updated: 2026-08-30
Code: https://github.com/CHATS-lab/ppi-eval
Terminology
Sources
- PPI++: Efficient Prediction-Powered Inference
- Power Analysis for Prediction-Powered Inference
- Efficient Inference for Noisy LLM-as-a-Judge Evaluation
- Noisy but Valid: Robust Statistical Evaluation of LLMs with Imperfect Judges
- Doubly-Robust LLM-as-a-Judge: Externally Valid Estimation with Imperfect Personas
- Predictions as Surrogates: Revisiting Surrogate Outcomes in the Age of AI
- Writing-Zero: Bridge the Gap Between Non-verifiable Tasks and Verifiable Rewards
- How to Correctly Report LLM-as-a-Judge Evaluations
- No Free Lunch: Non-Asymptotic Analysis of Prediction-Powered Inference
- A Note on the Prediction-Powered Bootstrap
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering