Two Calls, Two Moments, and the Vote-Accuracy Curve of Repeated LLM Inference

summary

Video file (mp4)

The gist

The gist: Two labeled calls identify mean accuracy and same-example correctness correlation, which in turn provide sharp distribution-free two-call intervals for any fixed majority-vote budget,

In short

The study uses two labeled calls to identify key moments of a latent success probability distribution from repeated LLM inferences. This allows for sharp, distribution-free intervals that separate stable errors from random call-level noise. It shows how majority voting can reduce many calls to a fixed budget while providing strong certification.

Key concepts

Latent Success Probability (q)
This is the unknown probability that a single LLM call on an example will be correct. The study assumes this probability follows some underlying distribution, which is what they are trying to characterize using the two calls.
Mean Latent Success Probability (µ)
The first labeled call identifies the average success rate (µ) of the latent probability distribution. This moment helps define the center of where the true success probabilities lie across all possible examples.
Same-Example Correctness Correlation (ρ)
This second moment captures how correlated the correctness is between different calls on the same example. It measures how much information one call provides about another, helping to quantify residual randomness after averaging.

Terminology used across episodes

This episode discusses

The paper

Two Calls, Two Moments, and the Vote-Accuracy Curve of Repeated LLM Inference · Read on arXiv

Yi Liu

York University

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "Two Calls, Two Moments, and the Vote-Accuracy Curve of Repeated LLM Inference".

Tom: The gist: Two labeled calls identify mean accuracy and same-example correctness correlation, which in turn provide sharp distribution-free two-call intervals for any fixed majority-vote budget,

Jane: First, who's behind it and why it matters.

Title and authors: Tom: Alright, let’s start with the title itself. "Two Calls, Two Moments, and the Vote-Accuracy Curve of Repeated LLM Inference." It’s a pretty descriptive name for a lot of math under the hood.

Jane: It points out that we're not just looking at one call accuracy; we're looking at how that repeats across different examples and how those repeats interact with each other.

Tom: Right. The authors are Yi Liu from York University, and they’ve done this work by focusing on the binary correctness layer of repeated inference under conditional-i.i.d. calls, which means we assume the examples are independent given some latent success probability Q, which is what they call the population law of example-level success probabilities (reference: <ref:2605.03379#pg1>).

Lu: It’s interesting because they connect this to how you get that vote-accuracy curve, which is V n = E P n(Q) for n calls, and they are saying we can identify the relevant parts of that curve with just two moments.

Meng: So, if I understand it right, instead of needing thousands of calls to map out the whole success pattern for a given prompt type, two carefully chosen labeled calls give us enough info to define the important boundaries.

Lalam: Right. It simplifies what used to be a huge validation problem into this much smaller moment problem.

The paper's summary: Tom: So, let’s get into the actual mechanism they propose because that’s where the payoff is. They show that one labeled call just tells you the mean success probability, mu, which is E Q.

Jane: That mu is easy to find, but it doesn't tell you much about how spread out or variable those successes are across different examples.

Tom: That’s where the second labeled call comes in; it identifies the second moment, nu, which is E Q2. And from those two values, they can calculate this same-example correctness correlation, rho = (nu - mu two)/(mu(one-mu)) (reference: <ref:2605.03379#pg2>).

Lu: That rho is the key piece of information because it tells you how much those repeated calls actually add independent evidence to each other, which separates stable errors from just random noise in the calls.

Meng: So, if I see a low correlation, that means the extra calls aren't helping us reliably because they are just introducing call-level randomness that we can’t average out with more prompting.

Lalam: That’s precisely what they are trying to distinguish from genuine improvements in accuracy that come from having more information.

Tom: And the result of this two-call identification is a sharp distribution-free two-call interval for any fixed majority-vote budget, which is huge because it means we can get exact bounds instead of fuzzy estimates.

The paper's improvements: Jane: What they suggest as an improvement is the massive technical reduction they make. They argue that you don't need to impose a full latent parametric model of Q to solve this problem.

Tom: That’s the big shift. They say that optimizing over all possible laws for Q using just those two moments boils down to a three-atom problem and an equivalent quadratic dual certificate (reference: <ref:2605.03379#pg2>).

Lu: That reduction is powerful because it means we can derive sharp endpoints for all finite budgets, and they even show the smallest deployable majority budget is n=one which means three calls (reference: <ref:2605.03379#pg1>).

Meng: Three calls as the minimum? That makes sense if the math allows it. So they get a closed-form three-vote interval with a width at most one/eight and they have this certified-improvement criterion that’s better than just guessing <ref:2605.03379#pg1,width at most 1/8, and>.

Lalam: It moves us away from needing to perform many repeated labeled calls per example to get that certification.

Tom: And they also introduce completions, like the maximum-entropy completion and the Latent-difficulty Gaussianprobit model, which help summarize the remaining ambiguity if we don't have those two exact moments.

Conclusion: Jane: So, wrapping up on this paper "Two Calls, Two Moments, and the Vote-Accuracy Curve of Repeated LLM Inference," what does it mean for us in practice? It seems like a major step toward making our repeated inference planning much more rigorous.

Tom: It’s about separating the stable error policies from those where call-level randomness can still be averaged away by majority voting, which is a huge clarification for how we evaluate AI systems.

Lu: The paper really boils down to this two-call statistic being able to distinguish those two things, and that empirical validation shows that three and five votes are empirically contained within the projected two-call regions (reference: <ref:2605.03379#pg1>).

Meng: From an engineering standpoint, it gives us a clear way to decide if we should invest compute into more calls or stop because the benefit is just noise.

Lalam: I think it helps guide future system design by giving us these sharp, distribution-free two-call certificates for any budget.

Tom: This paper provides a solid mathematical framework for moving from guessing the full latent law of success probabilities to using just two moments to get actionable, guaranteed bounds on our performance.

Jane: It’s definitely a piece of work that gives us much more confidence in how we plan complex interactions with LLMs, and it sets up some interesting directions for how we can use these moment problems to guide model selection.

More episodes

← Home