Correct Looks Better: Pairwise Comparisons Reveal Accuracy Rankings

arXiv:2606.09409 · cs.AI, cs.CL, cs.LG · Submitted 2026-06-08 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "Correct Looks Better".

Tom: Pairwise comparisons combined with aggregation methods like Elo have become central to evaluating generative models, yet concerns remain that they reward superficial stylistic cues or display judge biases.

Jane: First, who's behind it and why it matters.

Title and authors: Tom: So, to get into specifics about this paper, "Correct Looks Better: Pairwise Comparisons Reveal Accuracy Rankings," it’s essentially showing how using pairwise comparisons, like asking one model which of two answers it prefers, can actually give us a ranking that lines up surprisingly well with the actual ground truth accuracy scores when we have them.

Jane: That sounds like a big deal because traditionally, we often rely on direct evaluation methods to rank models, and this paper suggests that gathering those pairwise comparisons can recover a similar ordering to what you'd get from having human judges grade the answers.

Lu: What they’re demonstrating is that when you convert five well-known benchmarks into free-form questions and then use Elo-style rankings derived from these pairwise comparisons, the results show a Spearman correlation above zero point nine with accuracy rankings for MMLU Pro, GPQA Diamond, Simple QA, GSM8K, and BBH (Multitask).

Meng: A correlation above zero point nine is quite high; that suggests the method isn't just guessing; it’s actually capturing a lot of the same ordering you'd see if you had used labeled evaluations to rank them directly. I wonder how robust this is when we move past those specific benchmarks to more complex tasks.

Lalam: The fact that it works consistently across all five benchmarks really shows that the underlying mechanism—the preference aggregation—is quite stable, regardless of whether the task is multiple choice or free-form generation.

Tom: Exactly, and this consistency is what makes it so compelling; it suggests we can trust these preference rankings when ground truth data exists for comparison. It’s a solid foundation for how we assess model performance when direct accuracy metrics are hard to obtain.

The paper's summary: Jane: So, the main point of the paper is that they took five standard benchmarks and turned them into free-form questions, then used pairwise comparisons to rank models using an Elo-style system based on judge preferences. They found that these preference rankings strongly agree with accuracy rankings when ground truth is available for comparison.

Lu: The core finding is that Bradley-Terry rankings achieve high agreement with accuracy, specifically they report an Avg Spearman's Rho of zero point nine zero one nine ± zero point one zero and an Avg Kendall's Tau of zero point seven nine nine eight ± zero point one zero across those benchmarks, which consistently stays above zero point nine for the accuracy and Bradley-Terry scores together in Figure one.

Meng: That correlation data is pretty impressive, especially when you look at the setting where pairwise comparisons are most needed, which is what they call the evaluation frontier where judges are weak or far from solving the task correctly.

Tom: Right, and that's a crucial part of it; they found that in that specific regime—when the judge model performs poorly on its own—pairwise comparisons lead to substantially better alignment with accuracy than directly asking the judge about correctness. That really highlights the utility when evaluator quality is uncertain.

Lalam: It’s interesting because it shifts our focus from just trusting a single, potentially flawed judge to aggregating many pairwise judgments, which seems much more resilient when we don't have perfect supervision for every single output.

Jane: And they also looked into what affects these rankings in terms of style and bias, finding that these factors only have minor effects on the final rankings even though most judgments happen on pairs where both answers are correct or incorrect.

The paper's improvements: Tom: So, when we look at how the authors suggest we can make this better, they focus on understanding what drives preference signals in specific scenarios and how to refine the aggregation process itself.

Lu: They identified that on non-discriminative pairs—which is when both answers are either correct or incorrect—there’s a surprising amount of useful information present, specifically pointing to repetition after the final answer, or echo, as a causal driver of judge preference on those pairs.

Meng: Echo sounds like something we need to track in our logging systems; if a model repeats content after it thinks it's finished, that signals a failure mode that judges seem to pick up on even when the answers aren't perfectly discriminative.

Tom: And this is interesting because they found echo disappears on discriminative pairs where only one answer is correct, suggesting that when ground truth is available, the judge relies more on correctness than on these stylistic or coherence artifacts.

Jane: They also explored bias correction, testing if adjusting for style—like answer length or formatting—and self-preference biases improved the rankings, and they found that bias correction offered modest improvements on some of their benchmarks.

Lalam: Even with those modest improvements, the paper concludes that these biases have a limited impact on model ranking overall because almost fifty-eight percent of comparisons are non-discriminative, meaning judges still need other signals besides pure correctness.

Conclusion: Tom: So to wrap up this discussion on "Correct Looks Better: Pairwise Comparisons Reveal Accuracy Rankings," the main message is that using pairwise comparisons combined with methods like Bradley-Terry aggregation gives us rankings that are highly correlated with accuracy when ground truth is available, and it’s particularly useful in tricky evaluation situations.

Jane: It really suggests that these preference-based rankings can recover essentially the same ordering you would get from labeled evaluations, which is a solid piece of evidence for using pairwise comparisons as a robust tool.

Lu: The implication here is that we don't have to rely solely on perfect accuracy metrics; we can use this mechanism to get reliable relative performance estimates across several complex benchmarks simultaneously.

Meng: From an engineering standpoint, if we need a stable ranking quickly in an evaluation frontier where the judge is weak, this methodology seems like a much more dependable path than just relying on the direct output of that weak judge.

Lalam: I think what this paper really shows us is that quality isn't just about being factually correct; it involves judging coherence and flow in ways that pairwise comparison methods effectively capture, especially when we look at signals like echo.

Tom: Absolutely, and so while the study is scoped to benchmarks with ground truth answers, the results show a very strong alignment between preference-based rankings and those accuracy metrics. We’ll be keeping an eye on how this technique evolves as we build more complex generative systems.

Jane: It's definitely a paper that gives us another tool in our evaluation toolkit, so I think it’s worth checking out for anyone working on comparing these large models right now.

Lu: Indeed, and with the insights from "Correct Looks Better: Pairwise Comparisons Reveal Accuracy Rankings," we can start thinking about more sophisticated ways to measure generative performance beyond simple accuracy scores alone.

Meng: I'm eager to see how this translates into practical pipelines for iterative model refinement when we’re pushing the limits of what these systems can do.

Lalam: I just hope this research helps us build models that aren't just right, but are also coherent and engaging in their responses.

Max Planck Institute for Intelligent Systems

cs.AI, cs.CL, cs.LG

Submitted: 2026-06-08

Updated: 2026-06-08

Code: https://github.com/socialfoundations/correct-looks-better

Importance score: 91/100

The gist: Pairwise comparisons combined with aggregation methods like Elo have become central to evaluating generative models, yet concerns remain that they reward superficial stylistic cues or display judge

Key concepts

Pairwise Comparisons
This method involves generating answers from models and having a large language model judge the preference between two answers. This preference is used to create a ranking system, similar to Elo ratings, to estimate which model is better.
Bradley-Terry Model
This statistical model is used to aggregate the pairwise preferences into a single strength score for each generative model. It provides a maximum likelihood estimate of the models' true strengths based on all the comparison data collected.
Accuracy Alignment
This measures how well preference-based rankings align with rankings based on ground truth correctness. The research found that when ground truth is present, pairwise comparisons effectively recover the same ordering as accuracy-based labels.
Non-Discriminative Pairs
These are pairs of answers where both options are either correct or both are incorrect. The study found these pairs contain surprising amounts of useful information for judges, unlike discriminative pairs where only one answer is correct.

Terminology

Summary

Pairwise comparisons combined with aggregation methods like Elo have become central to evaluating generative models, yet concerns remain that they reward superficial stylistic cues or display judge biases. In a more positive turn, we show that model rankings from pairwise comparisons strongly agree with ground-truth-based accuracy rankings when such ground truth is available for comparison.

How it works

The study converts five well-known benchmarks into free-form generative evaluations and scores models using Elo-style rankings derived from pairwise comparisons. The core methodology involves:

  1. Generating answers to freeform questions from the models being evaluated.

  2. Collecting pairwise comparisons by randomly sampling pairs of model answers and prompting a large language model to express its binary preference over those pairs, asking which answer it is more confident in that it answers the question correctly.

  3. Aggregating these preferences using a Bradley-Terry model to estimate the strength of each model, resulting in an Elo-style ranking.

Key Findings on Accuracy Alignment

The research demonstrates that preference-based rankings align surprisingly well with accuracy rankings when ground truth is available. The results show:

Bradley-Terry rankings achieve high agreement with accuracy,

Spearman rank correlation above 0.9

This suggests that pairwise comparisons recover essentially the same ordering one would obtain from labeled evaluation. This alignment holds across five benchmarks: MMLU Pro, GPQA Diamond, Simple QA, GSM8K, and BBH (Multitask).

Robustness to Style and Bias

The study investigates the impact of stylistic cues and judge biases on model rankings. The findings indicate that these factors have only minor effects on the final rankings:

Style and judge bias have only minor effects on model rankings, despite most judgments occurring on pairs where both candidate answers are correct (or incorrect).

The researchers also find that repetition after the final answer (echo) is a causal driver of judge preference specifically on non-discriminative pairs.

Performance in Different Judge Regimes

The paper contrasts the performance of pairwise comparisons against a direct judge baseline, particularly in the evaluation frontier where judges are weak:

When the judge model performs poorly on the task, pairwise comparisons lead to substantially better alignment with accuracy than directly asking the judge about correctness.

This suggests that pairwise-preference rankings from weak judges remain well-aligned with the accuracy ranking, offering a robust alternative when evaluator quality is uncertain.

Signal in Non-Discriminative Pairs

The analysis of non-discriminative pairs (where both answers are either correct or incorrect) reveals useful information:

Non-discriminative pairs contain a surprising amount of useful information.

Specifically, the researchers identify echo as a causal driver of judge preference on these pairs. Echo is defined as a failure mode where a model fails to stop after its final answer and repeats content, such as duplicating phrases or re-generating the original question-answer template. This effect disappears on discriminative pairs where exactly one answer is correct, indicating that when ground truth is available, the judge relies on correctness rather than echo.

Evaluation Metrics and Aggregation

The study employs several metrics to quantify alignment:

Rank correlation We use Kendall’s Tau τ = C−D / C+D

Kendall’s distance KD = 1−τ2 is the fraction of model pairs that must be swapped to transform one ranking into another.

The Bradley-Terry model was chosen as the primary aggregation method because it provides the maximum likelihood estimate of model strengths given the full dataset, and it achieved the highest mean rank correlation across all tested benchmarks. The study also compared Bradley-Terry against Elo, WinRate, and TrueSkill, finding that Bradley-Terry generally achieved higher rank correlation.

Bias Correction Results

The researchers tested whether correcting for known biases—style (answer length, formatting), self-preference (model family bias), and a combination of both—led to improvements:

We find that bias correction offers modest improvements on some of our benchmarks.

However, the study concludes that these biases have limited impact on model ranking, noting that almost 58% of comparisons are non-discriminative, meaning judges must rely on other signals besides correctness. The most significant signal was found when restricting comparisons to only discriminative pairs (one answer correct, one incorrect).

Limitations and Future Directions

The study notes several limitations:

Discriminative tasks only. Our results are scoped to benchmarks with ground-truth answers.

The authors also acknowledge that MCQ and freeform rankings may differ, as they convert multiple-choice questions to freeform format for the study. Furthermore, the findings are "scoped to benchmarks with ground truth answers, which is what enables a controlled comparison with accuracy-based rankings.

Improvements for AI systems

Here are specific improvements for AI systems derived from this research:

  1. Improve model selection and ranking robustness by employing Elo-style pairwise comparison aggregation (Bradley-Terry) instead of relying solely on direct judge evaluations, especially in frontier evaluation settings where ground truth is unavailable or expensive. This allows for more stable relative rankings even when the judge is weak.

  2. Enhance the accuracy of relative model ordering by implementing bias correction mechanisms that explicitly account for style (answer length, formatting, etc.) and self-preference (model family). This prevents superficial cues from unduly influencing rankings, leading to more reliable performance metrics.

  3. Develop a robust evaluation pipeline capable of leveraging non-discriminative pairs (where both answers are correct or incorrect) as a source of signal. Instead of discarding these pairs, the system should analyze them for causal drivers like echo (repetition after the final answer).

  4. Implement an automatic echo detection system using LLMs to identify when models fail to stop generating by repeating sequences or new question-answer pairs. This detection serves as a proxy for quality/coherence and can be used to penalize or down-weight responses exhibiting this failure mode, particularly in non-discriminative settings where it is a known driver of preference.

  5. Utilize the Bradley-Terry model for strength estimation, as it consistently yields higher rank correlation with ground truth accuracy than simpler methods like Elo or WinRate across various benchmarks. This ensures that the derived model rankings are maximally aligned with true performance metrics when ground truth is available.

  6. For tasks where judge quality is uncertain (the evaluation frontier), use the Bradley-Terry ranking as the primary metric over a direct judge baseline, as it demonstrates superior robustness and less catastrophic failure in predicting accuracy compared to direct classification by a weak model.

Sources

Related papers