Sharp Limits for Honest Uncertainty in Hard-Budget Repeated Evaluation
cs.AI, cs.IT, math.IT
Submitted: 2026-09-24
Updated: 2026-09-24
Terminology
Sources
- Concentration inequalities for sampling without replacement
- Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
- An adaptation theory for nonparametric confidence intervals
- How Many Iterations to Jailbreak? Dynamic Budget Allocation for Multi-Turn LLM Evaluation
- Stratified Prediction-Powered Inference for Hybrid Language Model Evaluation
- Time-uniform, nonparametric, nonasymptotic confidence sequences
- How Many Tasks Are Enough for Agent Benchmark Decisions? A Replay Analysis of Public LLM Agent Benchmarks
- Efficient Prediction of Pass@k Scaling in Large Language Models
- Empirical Bernstein Bounds and Sample Variance Penalization
- Lazy ABC
- Sequential stratified inference for the mean
- Universal Inference
- Efficient Evaluation of LLM Performance with Statistical Guarantees
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection