Two Calls, Two Moments, and the Vote-Accuracy Curve of Repeated LLM Inference
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "Two Calls, Two Moments, and the Vote-Accuracy Curve of Repeated LLM Inference".
Tom: The gist: Two labeled calls identify mean accuracy and same-example correctness correlation, which in turn provide sharp distribution-free two-call intervals for any fixed majority-vote budget,
Jane: First, who's behind it and why it matters.
Title and authors: Tom: Alright, let’s start with the title itself. "Two Calls, Two Moments, and the Vote-Accuracy Curve of Repeated LLM Inference." It’s a pretty descriptive name for a lot of math under the hood.
Jane: It points out that we're not just looking at one call accuracy; we're looking at how that repeats across different examples and how those repeats interact with each other.
Tom: Right. The authors are Yi Liu from York University, and they’ve done this work by focusing on the binary correctness layer of repeated inference under conditional-i.i.d. calls, which means we assume the examples are independent given some latent success probability Q, which is what they call the population law of example-level success probabilities (reference: <ref:2605.03379#pg1>).
Lu: It’s interesting because they connect this to how you get that vote-accuracy curve, which is V n = E P n(Q) for n calls, and they are saying we can identify the relevant parts of that curve with just two moments.
Meng: So, if I understand it right, instead of needing thousands of calls to map out the whole success pattern for a given prompt type, two carefully chosen labeled calls give us enough info to define the important boundaries.
Lalam: Right. It simplifies what used to be a huge validation problem into this much smaller moment problem.
The paper's summary: Tom: So, let’s get into the actual mechanism they propose because that’s where the payoff is. They show that one labeled call just tells you the mean success probability, mu, which is E Q.
Jane: That mu is easy to find, but it doesn't tell you much about how spread out or variable those successes are across different examples.
Tom: That’s where the second labeled call comes in; it identifies the second moment, nu, which is E Q2. And from those two values, they can calculate this same-example correctness correlation, rho = (nu - mu two)/(mu(one-mu)) (reference: <ref:2605.03379#pg2>).
Lu: That rho is the key piece of information because it tells you how much those repeated calls actually add independent evidence to each other, which separates stable errors from just random noise in the calls.
Meng: So, if I see a low correlation, that means the extra calls aren't helping us reliably because they are just introducing call-level randomness that we can’t average out with more prompting.
Lalam: That’s precisely what they are trying to distinguish from genuine improvements in accuracy that come from having more information.
Tom: And the result of this two-call identification is a sharp distribution-free two-call interval for any fixed majority-vote budget, which is huge because it means we can get exact bounds instead of fuzzy estimates.
The paper's improvements: Jane: What they suggest as an improvement is the massive technical reduction they make. They argue that you don't need to impose a full latent parametric model of Q to solve this problem.
Tom: That’s the big shift. They say that optimizing over all possible laws for Q using just those two moments boils down to a three-atom problem and an equivalent quadratic dual certificate (reference: <ref:2605.03379#pg2>).
Lu: That reduction is powerful because it means we can derive sharp endpoints for all finite budgets, and they even show the smallest deployable majority budget is n=one which means three calls (reference: <ref:2605.03379#pg1>).
Meng: Three calls as the minimum? That makes sense if the math allows it. So they get a closed-form three-vote interval with a width at most one/eight and they have this certified-improvement criterion that’s better than just guessing <ref:2605.03379#pg1,width at most 1/8, and>.
Lalam: It moves us away from needing to perform many repeated labeled calls per example to get that certification.
Tom: And they also introduce completions, like the maximum-entropy completion and the Latent-difficulty Gaussianprobit model, which help summarize the remaining ambiguity if we don't have those two exact moments.
Conclusion: Jane: So, wrapping up on this paper "Two Calls, Two Moments, and the Vote-Accuracy Curve of Repeated LLM Inference," what does it mean for us in practice? It seems like a major step toward making our repeated inference planning much more rigorous.
Tom: It’s about separating the stable error policies from those where call-level randomness can still be averaged away by majority voting, which is a huge clarification for how we evaluate AI systems.
Lu: The paper really boils down to this two-call statistic being able to distinguish those two things, and that empirical validation shows that three and five votes are empirically contained within the projected two-call regions (reference: <ref:2605.03379#pg1>).
Meng: From an engineering standpoint, it gives us a clear way to decide if we should invest compute into more calls or stop because the benefit is just noise.
Lalam: I think it helps guide future system design by giving us these sharp, distribution-free two-call certificates for any budget.
Tom: This paper provides a solid mathematical framework for moving from guessing the full latent law of success probabilities to using just two moments to get actionable, guaranteed bounds on our performance.
Jane: It’s definitely a piece of work that gives us much more confidence in how we plan complex interactions with LLMs, and it sets up some interesting directions for how we can use these moment problems to guide model selection.
Yi Liu
York University
cs.LG, cs.CL
Submitted: 2026-05-05
Updated: 2026-10-03
Comments: 31 pages, 5 figures
Code: https://github.com/ollama/ollama
Project page: https://rajpurkar.github.io/SQuAD-explorer
License: http://creativecommons.org/licenses/by-sa/4.0/
Importance score: 91/100
The gist: The gist: Two labeled calls identify mean accuracy and same-example correctness correlation, which in turn provide sharp distribution-free two-call intervals for any fixed majority-vote budget,
Key concepts
- Latent Success Probability (q)
- This is the unknown probability that a single LLM call on an example will be correct. The study assumes this probability follows some underlying distribution, which is what they are trying to characterize using the two calls.
- Mean Latent Success Probability (µ)
- The first labeled call identifies the average success rate (µ) of the latent probability distribution. This moment helps define the center of where the true success probabilities lie across all possible examples.
- Same-Example Correctness Correlation (ρ)
- This second moment captures how correlated the correctness is between different calls on the same example. It measures how much information one call provides about another, helping to quantify residual randomness after averaging.
Terminology
Summary
The gist: Two labeled calls identify mean accuracy and same-example correctness correlation, which in turn provide sharp distribution-free two-call intervals for any fixed majority-vote budget, separating stable errors from recoverable call-level randomness.
How it works
The study isolates the binary correctness layer of repeated LLM inference where each call is reduced to a correctness bit. On an example with latent success probability q, a strict majority of 2n + 1 conditionally independent calls succeeds with probability Pn(q) = Pr[Bin(2n + 1, q) ≥ n + 1]The distinct vote-accuracy curve is Vn = EPn(Q) for n = 0, 1, 2,...>.
Two labeled calls are used to identify the first two moments of the latent correctness distribution: one call identifies the mean latent success probability µ = EQOne labeled call identifies only µ = EQ. A second labeled call identifies ν = EQ2 and the same-example correctness correlation ρ = (ν - µ2/µ(1 − µ))A second labeled call identifies ν = EQ2 and the same-example correctness correlation ρ = (ν - µ2/µ(1 − µ)). This pair table converts a potentially many-call validation problem into a two-call moment problemThis pair table converts a potentially many-call validation problem into a two-call moment problem.
Sharp Two-Call Certification
The central reduction is that this two-moment ambiguity class is computationally tractable without imposing a latent parametric modelThe central reduction is that this two-moment ambiguity class is computationally tractable without imposing a latent parametric model. For every fixed finite vote budget, optimizing over all laws on [0, 1] with moments (µ, ν) reduces to a three-atom problem and an equivalent quadratic dual certificateFor every fixed finite vote budget, optimizing over all laws on [0, 1] with moments (µ, ν) reduces to a three-atom problem and an equivalent quadratic dual certificate. This reduction allows for the derivation of sharp endpoints for all finite budgetsThis reduction allows for the derivation of sharp endpoints for all finite budgets. The smallest deployable majority budget is n = 1, i.e., three callsThe smallest deployable majority budget is n = 1, i.e., three calls. The closed-form three-vote interval is derived from this geometryThe closed-form three-vote interval is derived from this geometry.
Infinite-Vote Endpoint
The infinite-vote endpoint is the limit of the same majority rule as the number of conditionally independent calls goes to infinityThe infinite-vote endpoint is the limit of the same majority rule as the number of conditionally independent calls goes to infinity. This endpoint is not a new voting rule but a threshold functional of the latent lawThis endpoint is not a new voting rule but a threshold functional of the latent law. It equals the latent mass above the threshold q = 1/2 plus half the mass exactly at the thresholdIt equals the latent mass above the threshold q = 1/2 plus half the mass exactly at the threshold.
Model-Based Completions
Two completions are added to summarize the remaining ambiguity: maximum-entropy and Latent-difficulty Gaussianprobit (LDGP)Two completions are added to summarize the remaining ambiguity: maximum-entropy and Latent-difficulty Gaussianprobit (LDGP). The maximum-entropy completion is the least-informative law on [0, 1] with the observed momentsThe maximum-entropy completion is the least-informative law on [0, 1] with the observed moments. The LDGP completion uses the established normal-ogive/probit item-response difficulty model for binary correctnessThe LDGP completion uses the established normal-ogive/probit item-response difficulty model for binary correctness.
Empirical Validation
Experiments on LLM calls over QNLI and QQP show that empirical three- and five-vote accuracies are contained in the projected two-call regionsExperiments on LLM calls over QNLI and QQP show that empirical three- and five-vote accuracies are contained in the projected two-call regions. The results show that empirical three- and five-vote accuracies are contained in the projected two-call regions. Lower same-example correlation corresponds to residual call-level randomness that voting can average awayLower same-example correlation corresponds to residual call-level randomness that voting can average away.
Discussion
The paper separates repeated-inference planning into an identifiable two-call certification component and a parametric completion componentThe paper separates repeated-inference planning into an identifiable two-call certification component and a parametric completion component. The two-call statistic distinguishes stable-error policies from policies whose residual call-level randomness can still be converted into finite-vote accuracy by majority votingThe two-call statistic distinguishes stable-error policies from policies whose residual call-level randomness can still be converted into finite-vote accuracy by majority voting. The empirical results express the same separation operationallyThe empirical results express the same separation operationally.
Computation Details
The experiments use GLUE QNLI and QQP with three local model families and five repeated calls per exampleThe experiments use GLUE QNLI and QQP with three local model families and five repeated calls per example.
Improvements for AI systems
-
Bold header: One-Call Accuracy Limitation Overcome by Two-Call Moment Identification. The system can distinguish between
stable errors from recoverable call-level randomness
by identifying thesame-example correctness correlation that separates stable errors from recoverable call-level randomness,
which allows for a more informed decision on whether to proceed with repeated inference planning. -
Bold header: Distribution-Free Certification for Finite Vote Budgets. The system can obtain
sharp distribution-free two-call certificates for every finite majority-vote budget,
providing exact bounds, such as theclosed form three-vote interval
and acertified-improvement criterion,
rather than relying on discretized or parametric approximations. -
Bold header: Optimal Planning via Three-Atom Geometry. The system can optimize planning objectives by leveraging the key technical reduction that
the infinite-dimensional moment problem has three-atom extremizers,
allowing it to derive the sharp endpoints for any finite budget through aquadratic dual certificate.
-
Bold header: Model Selection Guided by Latent Difficulty Models. By using point completions like Latent-difficulty Gaussian-probit (LDGP), the system can select a specific latent law that matches observed two-call moments, resulting in a
whole vote curve
and potentially more accurate performance than simple interval midpoints. -
Bold header: Overtaking Policies Identification. The system can identify scenarios where
lower one-call accuracy may beat a stronger one-call baseline after voting when it has lower same-example correlation,
enabling the deployment of randomized policies that exploit residual call-level randomness for significant vote gains.
Abstract
Repeated sampling can improve LLM accuracy, and quantifying the gains from additional calls is essential for allocating test-time compute. We study binary decisions, a fundamental setting where repeated answers to the same question are aggregated by majority vote. We show that two independently sampled responses per example in a validation set with known answers constrain the latent distribution of example-level success probabilities to a class consistent with the paired outcomes. Optimizing over this class yields sharp accuracy and gain bounds at every finite voting budget. A shared large-sample confidence region accounts for validation uncertainty. We also obtain sharp infinite-vote bounds and moment-matched forecasts of the vote-accuracy curve. On QNLI and QQP, our method successfully distinguishes settings with voting gains above a target margin from those with little room for improvement.
Sources
- Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
- Evaluating Large Language Models Trained on Code
- Estimating the Self-Consistency of LLMs
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks