The Judge Knows When It Knows: Calibrated Abstention for LLM-Based A/B-Test Prediction
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "The Judge Knows When It Knows: Calibrated Abstention for LLM-Based A/B-Test Prediction".
Jane: The paper was written by Tyler Dooskin and the Squoosh Technical Staff from Squoosh.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back to the arXiv review show, everyone. I'm Tom, and today we're digging into a paper that has one of the most honest titles I've seen in a long time: "The Judge Knows When It Knows: Calibrated Abstention for LLM-Based A/B-Test Prediction."
Jane: And I'm Jane. Tom, that title is doing so much work. It's not claiming the judge is always right. It's claiming the judge knows when it knows. That's a completely different product.
Tom: Exactly. And it's from the team at Squoosh, which is an AI startup working on conversion optimization. They basically asked a very simple question: can a multimodal LLM look at two screenshots of a web page and predict which one will win a real A/B test?
Jane: And the answer, spoiler alert, is mostly no. But the way they get to that answer is what makes this paper special. They ran six weeks of pre-registered experiments on real conversion tests.
Tom: Pre-registered. That's the key phrase. They locked their hypotheses and their decision rules before they even collected the data. That's not how most commercial AI research works.
Jane: Right. And their headline finding is that unconditional winner prediction doesn't clear what they call the honesty bar. On three hundred thirty real A/B tests, their judge gets a Cohen's kappa of zero point one four one. That's detectably above chance, but it's weak.
Tom: And here's the kicker. When they look only at the statistically significant tests — the ones where we actually know who won — the evidence becomes inconclusive. Kappa drops to zero point one zero eight, and the confidence interval includes zero.
Jane: So the model is basically guessing on the tests where we have trustworthy answers. But it's agreeing with the unreliable labels more than the reliable ones. That's the signature of a shared prior, not prediction.
Tom: A shared prior. Can you unpack that for our listeners?
Jane: Sure. The corpus they used comes from a CRO agency's case library. That library is curated by humans who have their own folk wisdom about what should win. The LLM was trained on the same folk wisdom. So when the model agrees with the curator on a marginal test, they're not predicting the outcome. They're both just expressing the same prior belief.
Tom: And that's why the title is so important. The judge doesn't always know. But the paper argues the judge can learn to know when it knows, and abstain otherwise.
Jane: That's the calibrated abstention part. Instead of forcing an answer on every test, the judge only makes a call when its internal panel of votes is confident enough. And that's where the signal actually appears.
Tom: So the title is a promise, and the paper delivers on it. But it's a much more modest promise than what most vendors are selling.
Jane: And that's exactly why this paper matters. It's a blueprint for how to do honest AI evaluation in a commercial setting. Stick around, because next we're going to dig into the actual results and what the judge can and cannot do.
Summary: Tom: Welcome back. We're still on "The Judge Knows When It Knows," and Jane just walked us through the setup. Now let's talk about what they actually found when they started testing the obvious fixes.
Jane: Right. So you'd think, okay, the model is weak. Let's just use a bigger model. Let's use a better prompt. Let's give it better screenshots. The paper tests all of those, and they all fail.
Tom: And they fail in a really clean way. They compared Gemini three Flash against Gemini three point one Pro on the same one hundred fifty-nine tests. The difference in kappa was minus zero point zero three seven, with a confidence interval that lies entirely inside their pre-declared equivalence bound.
Jane: So they're not just saying "we didn't find a difference." They're saying "we can statistically prove these two models are equivalent on this task." That's a powered equivalence claim, not an absence-of-evidence shrug.
Tom: And the Pro model costs two point eight times more and is slower. So model capacity is not the bottleneck. That's a confirmed negative result.
Jane: Then they tried prompt redesign. They had a calibration-targeted rewrite that reduced the model's tendency to over-predict the control arm, but it still failed the kappa non-inferiority test. And they tried stimulus fidelity — real screenshots versus reconstructed pages. Parity. They tried change-type priors — like "this category of change usually wins." That predicted held-out outcomes at kappa zero point zero one five, which is basically zero.
Tom: So every lever you'd naturally pull to make this work just doesn't move the needle. And that's actually a really important negative result for the industry.
Jane: But then they found the one thing that does work. They have this panel of sixteen to twenty votes per test, and they gate on the margin between votes for each arm. If the margin is big enough, the judge makes a call. Otherwise, it abstains.
Tom: And at a margin of zero point six, the judge calls forty-nine percent of tests and abstains on the rest. On the significant-only labels, that called subset reaches kappa zero point three one one, with a confidence interval that excludes zero.
Jane: That's the only cell in the entire research program where a significant-only confidence interval excludes zero. The judge's confident calls carry real signal, even on trustworthy labels.
Tom: But here's the honest part. That operating point is exploratory. They pre-registered a held-out confirmation at the stricter unanimity gate, and it came back inconclusive. Kappa minus zero point one zero zero on thirty-three calls.
Jane: And that's the moment where this paper really earns its credibility. They had an in-sample estimate of kappa zero point three five, and the held-out test didn't directionally replicate. They report it as exactly that — an underpowered null that upgrades nothing and refutes nothing.
Tom: And they're running a powered confirmation on partner data right now. That's the designed escape from the single-corpus problem.
Jane: But the most fascinating finding to me is the reliability-versus-validity result. The held-out run's unanimous calls agreed with an independent earlier run on thirty-one out of thirty-one overlapping tests. Perfect test-retest reliability. And yet the judge scored kappa minus zero point one zero zero against truth.
Tom: So the judge is perfectly consistent and completely wrong at the same time.
Jane: Exactly. Unanimity in an LLM panel is one deterministic prior expressed many times, not independent evidence accumulating. That's a warning for anyone building LLM judge ensembles.
Tom: And that's the deep cut of this paper. Next segment, we're going to bring in Lu and Meng to talk about what this means for the practical deployment of these systems.
Improvements: Tom: Welcome back. We're deep in "The Judge Knows When It Knows," and I want to bring in Lu and Meng now, because this paper has some really concrete implications for how you'd actually build and deploy this kind of system.
Jane: And I think the most important improvement the paper suggests is the shift from always-answer to confident-call-or-abstain. That's not just a tweak. That's a different product philosophy.
Lu: Right, Jane. And I think the deepest result here is the measurement of the shared prior. They showed that two judges differing in model or prompt agree with each other at kappa zero point seven four to zero point eight eight, while each agrees with real outcomes at only about zero point two. That's a three-to-four-fold gap.
Tom: And that gap is the whole story, isn't it? The judges are not sampling independent evidence. They're all expressing the same training prior.
Lu: Exactly. And they quantify it. A sixteen-vote panel carries about two effective independent votes. One-way ICC of zero point three four to zero point five one. So when you see unanimous agreement in an LLM panel, you should not read that as accumulating evidence. You should read it as one deterministic prior expressed many times.
Meng: And that's the thing that keeps me up at night as an engineer. Because reproducibility is the thing we usually chase. If I can make the model produce the same answer every time, I feel good. This paper shows that perfect reproducibility can coexist with zero validity.
Jane: That's the reliability-versus-validity trap. And they caught it in their own held-out run. thirty-one out of thirty-one identical calls across runs, and kappa minus zero point one zero zero against truth.
Meng: So what do we actually ship? The paper is pretty clear. You ship the gate. You ship abstention as a first-class behavior. The judge only makes a call when the vote margin clears the threshold, and otherwise it says "I don't know."
Tom: And the economics work out too. They did vote-subsampling and found a nine-vote panel preserves the gated operating point at roughly half the inference cost of the full panel. And the proxy tier runs at about two cents per evaluation versus about a dollar twenty-five for a browser-agent tier.
Meng: That's the part that makes this deployable. You can screen an entire experiment backlog for less than the cost of running one shopper through one arm of one live test.
Lu: And there's a deeper implication here. They measured the human ceiling too. Two hundred plus CRO professionals averaged twenty-nine percent on eight real tests — chance. And their own fifteen-expert baseline reproduced the shared-prior collapse directly. Experts agree with each other at kappa zero point five three but score at chance against real outcomes.
Jane: So the human experts have the same problem as the LLM. They share folk wisdom about what should win, and that folk wisdom is uncorrelated with what actually wins.
Lu: That's the paper's thesis in its strongest form. On this task, consensus — human or model — is reproducible, persuasive, and not evidence.
Tom: And that reframes what the product is. It's not a win-rate oracle. It's a calibrated screening instrument that knows which minority of cases are callable at all.
Meng: And the paper is very disciplined about what it does not claim. No predicted conversion-lift percentages. No mobile-validated accuracy, because no mobile-rendered labeled corpus exists. Every externally quoted number carries an evidence tier.
Jane: That claims ledger in Appendix A is a model for the whole industry. A number quoted without its tier tag is out of policy. That's how you build trust in a field drowning in vendor hype.
Tom: And that's the improvement the paper suggests, really. Not a better model, but a better discipline. Let's wrap this up in the conclusion.
Conclusion: Tom: And we're back for the final segment on "The Judge Knows When It Knows: Calibrated Abstention for LLM-Based A/B-Test Prediction." Jane, give us the one-paragraph version.
Jane: This paper asks whether an LLM can predict A/B test winners from screenshots, and the answer is: mostly no, but the exceptions are identifiable in advance. Unconditional prediction is statistically unprovable on trustworthy labels. But gating on panel confidence isolates a subset of calls with real signal, and the judge abstains on everything else.
Tom: And the paper's real contribution is the discipline. Pre-registration, locked gates, negative results published, every claim tiered. They even caught their own best results failing their own audits.
Jane: That's the part I keep coming back to. They had an in-sample gated estimate of kappa zero point three five. Their held-out confirmation came back inconclusive. And they reported it as exactly that. A measurement system that cannot kill its owner's claims is advertising, not measurement.
Tom: And the human baseline makes the result more meaningful. Experts are at chance on this task too. So a judge that knows when it knows, even at kappa zero point three one, is a screening instrument worth having.
Meng: And the engineering path is clear. Nine-vote panels at half the cost, two cents per evaluation, abstention as a default behavior. That's deployable today.
Lu: And the warning for the field is the shared-prior collapse. Inter-judge agreement is not evidence. Consensus among correlated judges is reproducible, persuasive, and not evidence. That applies to LLM panels and to human expert panels alike.
Tom: So we're saying goodbye to this paper, and it's a rare one. It's not a hype paper. It's a calibration paper. It tells you exactly what the tool can do, exactly what it cannot do, and exactly how to tell the difference.
Jane: And that's the standard we should hold every AI product to. Thanks for listening, everyone. Next up, we've got a paper on generative agents for user simulation, and we'll see if it holds up to the same scrutiny.
Tom: See you then.
Tyler Dooskin, the Squoosh Technical Staff
Squoosh
cs.HC, cs.CL, stat.AP
Submitted: 2026-07-02
Comments: 15 pages. Pre-registered experimental program with a public, tiered claims ledger; includes powered negative results, a label-validity audit, a cross-judge shared-prior measurement (n_eff ~ 2 of 16 votes), and a first-party 15-expert human baseline. Pre-registrations, statistical harness, human responses, and the full experiment ledger are released
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 74/100
Terminology
Summary
Summary
This paper reports on a six-week, pre-registered experimental program investigating whether a multimodal LLM can predict the winner of real A/B tests from screenshots alone. The central finding is that unconditional winner prediction does not clear the honesty bar
: on 330 real A/B tests, a Gemini 3 Flash judge attains Cohen’s κ = 0.141 [0.034, 0.248] — detectably above chance, but on the trustworthy (statistically significant) half of the labels the evidence is inconclusive (κ = 0.108 [−0.049, 0.264]).
The paper establishes that 44% of the ground-truth
labels in the primary corpus — the curated case library of a leading CRO agency — come from non-significant tests, and that the judge agrees more with those unreliable labels than with reliable ones, which is the signature of a shared prior between label curator and model, not of prediction.
Every standard improvement lever fails: a 2.8× more expensive frontier model is statistically equivalent to Flash as the judge (∆κ = −0.037 [−0.152, +0.076], within a ±0.20 TOST bound, n=159 paired); prompt-mechanism redesign, stimulus fidelity, and change-type priors all fail their pre-registered gates.
The judge’s confident calls are different. Gating predictions on internal panel agreement concentrates real signal: at a vote-margin ≥ 0.6 gate the judge calls 49% of tests and abstains on the rest, and on significant-only labels the called subset reaches κ = 0.311 [0.025, 0.555] — "the only operating point that survives our own label-validity audit (exploratory tier; a pre-registered held-out confirmation at the stricter unanimity point returned κ = −0.100 [−0.200, 0.000] on an underpowered 33-call cohort, which by our locked decision rule upgrades nothing and kills nothing; a powered confirmation on partner data is pre-registered and in motion)."
The abstention behavior itself is robust: coverage replicates out-of-sample (27.8% → 30.3%), and the gated calls are perfectly reproducible across independent runs (31/31) — which the paper shows is exactly why reproducibility must never be mistaken for validity.
The mechanism is measurable: judges differing in model or prompt agree with each other at κ = 0.74–0.88 while each agrees with outcomes at only ∼ 0.2, and a 16-vote panel carries ≈ 2 effective independent votes.
The economics work: vote-subsampling shows a 9-vote panel preserves the gated operating point (κ = 0.264–0.271 at 46% coverage) at roughly half the inference cost of the full panel, and the proxy tier runs at ∼ 0.02 per evaluation versus ∼ 1.25 for a browser-agent tier.
The deliverable this evidence supports is "not a win-rate oracle. It is a calibrated screening instrument: confident-call-or-abstain winner reads, a grounded mobile-UX audit whose findings pass an adversarial verification pass (12 of 13 verified, 1 auto-rejected), and a published claims ledger in which every number carries an evidence tier."
For calibration of expectations: in the best published human benchmark, 200+ CRO professionals averaged 29% on 8 real tests — indistinguishable from chance; and a fresh 15-expert baseline on the authors' own instrument reproduces the shared-prior collapse directly — experts agree with each other (inter-rater κ = 0.53) but score at chance against the real outcomes (κ ≈ 0), with neff ≈ 1.8 effective independent experts of 15.
The paper releases its pre-registrations, locked gates, negative results, statistical harness, human responses, and the full experiment ledger.
Improvements for AI systems
Based on this paper, here are the specific improvements I can make to AI systems, and what the improved system can do:
Improvement: Replace unconditional prediction with a vote-margin gate. The system computes a panel of 16–20 counterbalanced votes per input, calculates the margin between the top two answers, and abstains (returns I don't know
) when the margin falls below a pre-declared threshold (e.g., margin ≥ 0.6 or unanimity).
What the improved system can do: On A/B-test winner prediction, the system now achieves κ = 0.311 [0.025, 0.555] on significant-only labels at 49% coverage, versus κ = 0.108 [−0.049, 0.264] without the gate. It refuses to answer on the 50% of inputs where it has no measurable skill, rather than emitting a confident guess that is wrong.
Improvement: Before reporting any accuracy or agreement metric, the system recomputes the statistical significance of each ground-truth label from its raw per-arm counts (sample sizes, successes). It then reports performance separately for significant-only labels versus all labels, and flags any case where the model agrees more with non-significant labels than significant ones (the shared-prior artifact signature).
Improvement: Before treating panel agreement or multi-agent consensus as evidence, the system computes two quantities: (a) inter-judge agreement (e.g., Cohen's κ between two independently configured judges differing in model or prompt), and (b) each judge's agreement with ground truth. It also computes the effective number of independent votes (neff) via one-way ICC. If inter-judge κ is high (e.g., 0.74–0.88) but each judge's κ vs. truth is low (e.g., 0.2), the system flags that consensus is a shared prior, not independent evidence.
Improvement: Replace raw accuracy with chance-corrected Cohen's κ as the primary gate metric in all evaluation and production settings. The system computes κ with percentile bootstrap CIs (B = 10,000, fixed seed) and uses gate-on-resample (re-applying the gate inside every bootstrap replicate) to propagate coverage uncertainty.
Improvement: For every prediction task, the system runs at least three trivial baselines alongside the main model: always-predict-the-majority-class, pixel-difference counters, and DOM-feature heuristics. The system reports the model's κ relative to these baselines and refuses to claim skill if the model does not exceed them on significant-only labels.
Improvement: Before any evaluation or deployment decision, the system commits to: (a) a locked hypothesis, (b) a pre-declared decision rule (e.g., κ ≥ 0.2 on significant-only labels with a minimum coverage), (c) a power analysis, and (d) a verdict space that includes INCONCLUSIVE (underpowered null upgrades nothing, refutes nothing). Prediction files and analysis outputs are locked before data collection.
Improvement: The system measures inter-rater agreement (e.g., pairwise κ, ICC) and effective independent raters (neff) for any panel — human or LLM. If inter-rater κ is high (e.g., 0.53) but each rater's κ vs. truth is near zero, the system flags the panel as consensus without validity
and refuses to treat majority vote as evidence.
Improvement: The system audits its own high-confidence calls (e.g., unanimous votes) against ground truth, computing κ separately for that slice. It also compares the slice's accuracy against the trivial baseline on the same slice.
Improvement: The system uses vote-subsampling (hypergeometric, seeded) to determine the minimum panel size that preserves the gated operating point within its uncertainty band. It then selects the cheapest panel that meets the precision requirement.
Improvement: Every quantitative claim the system emits carries an evidence tier (CONFIRMED, REPLICATED-DIRECTIONAL, EXPLORATORY, INCONCLUSIVE, REFUTED, DEMONSTRATED). The system refuses to emit a claim without its tier tag, and refuses to emit any conversion-lift percentage at all (standing rule).
-
Predict A/B-test winners with calibrated abstention — call a winner only when a vote-margin gate is cleared, abstain otherwise, achieving κ = 0.311 on trustworthy labels at 49% coverage.
-
Refuse to answer when it lacks skill — instead of emitting a confident guess on the 50% of inputs where it has no measurable signal.
-
Detect and report label-validity artifacts — stratify by statistical significance, flag shared-prior agreement, and never report raw accuracy as evidence.
-
Distinguish reproducibility from validity — measure inter-judge agreement vs. ground-truth agreement, and warn when consensus is a shared prior.
-
Optimize panel cost — use the smallest panel that preserves precision, cutting inference cost by 50–75%.
-
Enforce evidence tiers — every claim carries a locked tier, and unconfirmed claims are explicitly flagged as such.
-
Catch its own failures — via pre-registered gates, trivial baselines, and adversarial verification, the system can kill its own best results before they cause harm.
Abstract
Can a multimodal LLM predict which version of a web page will win a real A/B test from screenshots alone? We report the most complete answer we are aware of, from six weeks of pre-registered experiments on real conversion tests: mostly no -- and the exceptions are identifiable in advance. On 330 real A/B tests a Gemini 3 Flash judge reaches Cohen's kappa = 0.14, but on the trustworthy (statistically significant) half of the labels the evidence is inconclusive (kappa = 0.11, CI includes zero). We show that 44% of the "ground-truth" labels in a leading CRO agency's catalog come from non-significant tests, and that the judge agrees more with the unreliable labels than the reliable ones -- a shared prior between labeler and model, not prediction. Every standard improvement lever (a 2.8x more expensive frontier model, prompt redesign, stimulus fidelity, change-type priors) fails its pre-registered gate. The judge's confident calls are different: a vote-margin gate isolates a subset (49% coverage) reaching kappa = 0.31 on significant labels. We measure the mechanism directly -- judges differing in model or prompt agree with each other at kappa = 0.74-0.88 while agreeing with real outcomes at only 0.2, so a 16-vote panel carries about 2 effective independent votes -- and we reproduce it in humans: 15 CRO experts agree with each other (inter-rater kappa = 0.53) but score at chance against real outcomes (kappa 0). Consensus, human or model, is reproducible, persuasive, and not evidence. We release our pre-registrations, locked gates, negative results, statistical harness, human responses, and a claims ledger in which every number carries an evidence tier.
Sources
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
- G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment
- Large Language Models are not Fair Evaluators
- Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models
- Selective Classification for Deep Neural Networks
- Selective Question Answering under Domain Shift
- On Calibration of Modern Neural Networks
- Generative Agents: Interactive Simulacra of Human Behavior
- UXAgent: An LLM Agent-Based Usability Testing Framework for Web Design
Related papers
- EduGage: A Multimodal Dataset and Benchmark for Sensor-Based Momentary Assessment of Engagement in Self-Guided Video Learning
- EvoDesign: Agentic Editable Diagram Creation via Design Expertise Evolution
- HAGI++: Head-Assisted Gaze Imputation and Generation
- Linking Behaviour and Perception to Evaluate Meaningful Human Control over Partially Automated Driving
- Review of Explainable Decision Support and Adaptive Human-Machine Interfaces for Automation Transparency in Maritime Autonomous Surface Ships
- Towards Cognitive Process-Aware Proactive Writing Support