The First Token Is Not the Verdict: Hidden Costs of Reading LLM Judges Without Generating

arXiv:2610.00054 · cs.CL, cs.AI · Submitted 2026-09-04 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "The First Token Is Not the Verdict".

Jane: Reading an LLM judge’s verdict from its first generated token reveals that this cheap readout distorts position bias in one direction,

Tom: First, who's behind it and why it matters.

Paper summary: Tom: So, we're diving into a paper called "The First Token Is Not the Verdict: Hidden Costs of Reading LLM Judges Without Generating." Basically, the authors are looking at how easy it is to read an AI judge's decision just by looking at the very first piece of text it spits out.

Jane: That’s right, Tom. The core idea here is that this cheap way of getting a verdict—reading the logits from that initial token—isn't actually accurate; it actually introduces a sort of bias into how we measure things, meaning those figures end up acting more like an upper bound than a true measure of judgment.

Lu: From my perspective as someone who works on the architecture side, I think this highlights a fundamental disconnect in how we are trusting these initial outputs without seeing the full reasoning process behind them.

Meng: I’m curious about the practical impact here; if these readouts are misleading, how does that change how we deploy or trust AI systems in real-world applications?

Lalam: I see this as an opportunity for us to refine how we evaluate performance, making sure our internal metrics reflect reality rather than a shortcut.

Tom: Exactly. The paper points out that judges don't always start with a verdict token, which is where things get interesting because forcing a read on those non-verdict starts often returns whatever response was shown first instead of actually capturing the judgment.

Jane: It means there’s a significant difference between reading the first token directly and waiting for the full generation to see what happened, and this difference shows up very strongly in certain conditions.

Lu: The paper quantifies this divergence quite clearly, showing that when judges don't lead with a verdict token, the forced read flips on eighty-nine point seven percent of those pairs when the responses are swapped, compared to only forty-seven point five percent of them reading after generation.

Meng: That eighty-nine point seven percent flip rate is substantial; if we rely on the cheap readout, we're systematically overstating position bias by forty-two points, which is a lot of distortion for what’s measured.

Lalam: That distortion moves position bias by forty-two points while only moving judge accuracy by less than one point in seven of ten conditions, so the paper suggests this readout misleads those who audit the judge rather than those who actually use it.

Tom: It really puts things into perspective, doesn't it? We have a shortcut for getting an initial idea, but that shortcut is flawed when you're trying to measure the actual quality of the judgment itself.

Jane: And this leads us nicely into what the authors call position bias measurement and protocols, which they set up by comparing two different ways to get a verdict.

Paper summary: Lu: They are essentially contrasting a protocol where you generate a response and then parse the verdict from that text, versus one that reads the verdict directly from the logits of the first generated token.

Meng: From an engineering standpoint, it makes sense why this is cheaper because it avoids running a full generation just to check a single token, which is what constrained decoding and likelihood scoring evaluations do by construction.

Lalam: That's the trade-off they’re highlighting: the readout is cheap and requires no generation, but it comes with this inherent distortion because it doesn't always capture a true verdict token.

Tom: And the paper focuses specifically on what they call the forced read, which takes the higher of two verdict-token logits at that first generated position, and this specific property is what makes this paper so important.

Jane: It shows how different methods of extracting information from an AI output can yield fundamentally different results when you’re trying to assess a judgment.

Lu: The authors explore how existing practices handle non-verdict first tokens in three ways: the forced read, restricted-likelihood harnesses, and generating and then retrying or letting the judge continue.

Meng: I wonder if this forces a different kind of decision on the system; forcing a read on those pairs that don't commit seems to return whatever response was shown first, which is not a judgment at all.

Lalam: That’s the core issue with relying solely on that first token when the model isn't committed to a decision right away; it just defaults to the easiest path rather than assessing quality.

Tom: It really drives home that this readout is unreliable if you aren't careful about what you’re measuring, because it moves position bias by forty-two points while affecting judge accuracy by less than one point in seven conditions.

Jane: So, when we look at the results, the paper shows that this divergence between protocols is most noticeable when a judge did not lead with a verdict token.

Lu: The study pooled over nine hundred twenty-four pairs where a judge didn't commit and found that the forced read flips on eighty-nine point seven percent of those pairs when you swap the responses, which contrasts sharply with only forty-seven point five percent read after generation.

Meng: That massive difference in flip rates is what makes this distortion so pronounced, and it shows that the reliability of this cheap readout depends entirely on the model's behavior at that first step.

Lalam: It suggests a kind of adaptive computation might be necessary, where you only generate for those orders where the judge didn't commit, which is a standard practice but then escalating that to a stronger model under a calibrated threshold.

Paper summary: Tom: That adaptive approach sounds like a reasonable way to mitigate the risk, even though the paper says fidelity isn't perfect when you use those readouts.

Jane: And it really brings us to the recommendations for disclosure that the authors give, which are quite practical for anyone deploying these judge systems.

Lu: They suggest reporting three specific things in judge deployment disclosure: first, the compliance rate, which costs one forward pass per pair and can vary from zero point five zero six to one point zero zero zero across conditions.

Meng: And second, which readout produced the verdict because figures from constrained decoding and generate-and-parse aren't the same quantity when compliance is incomplete.

Lalam: They also suggest treating the answer format as an experimental variable, noting that changing just two characters can move compliance by seventeen points here.

Tom: So, the key message is that while reading the first token is a convenient place to check for a verdict, it’s also an unreliable one because there's up to forty-nine percent of pairs where it isn't actually a verdict at all.

Jane: It’s about being transparent about which readout you are using and acknowledging that the results you get from that cheap method might be inflated upper bounds.

Lu: This research contributes to understanding the nuances of how we evaluate LLM outputs, showing that even seemingly simple extra steps have complex effects on measurement fidelity.

Meng: I think this means for practical engineering, we need to be very precise about our evaluation pipelines and not just rely on the fastest method available.

Lalam: And from a cultural standpoint, understanding these measurement pitfalls helps us build more robust and trustworthy systems that reflect real judgment rather than just superficial data points.

Tom: That’s the gist of "The First Token Is Not the Verdict: Hidden Costs of Reading LLM Judges Without Generating." We've seen how this cheap readout distorts position bias, and it clearly shows that relying on it without generating is misleading for anyone trying to audit or use those judge results accurately.

Jane: It’s a crucial reminder that simplicity in measurement can sometimes lead to significant errors if you don't account for the underlying mechanism of what you are observing.

Lu: The implications for the broader field are that we need to be much more careful about how we define and measure "judgment" when working with these complex models.

Meng: I just hope this pushes us toward more rigorous evaluation standards across the board, because precision in measurement matters when you're building something that impacts people’s lives.

Lalam: It definitely gives us a stronger framework for building AI where we can be honest about what we know and what we don't yet understand about the model's decision process.

Conclusion: Tom: So, we've been diving deep into the paper "The First Token Is Not the Verdict: Hidden Costs of Reading LLM Judges Without Generating," and now we're getting to wrap up what all this means for us.

Jane: It boils down to how cheap it is to get a verdict from an AI judge just by looking at its first word, and the authors show that this shortcut actually distorts our understanding of the judge's actual decision-making process.

Lu: I think the core insight is that this quick readout isn't always what it seems because judges don't always start with a clear verdict token, which messes up how we measure things when we use it.

Meng: From an engineering standpoint, this means we can no longer just trust that first token as a reliable indicator of judgment quality without running the full process sometimes.

Lalam: It points toward a need for more rigorous evaluation methods because the way you extract information from an AI output has real, measurable consequences on what you measure.

Tom: Exactly, so when we look at the authors and what they've done, they're showing us that reading that first token is convenient but carries a hidden cost in terms of accuracy for auditing systems.

Jane: The authors are making a very simple point: if you rely on reading just the initial token without generating the rest, you’re looking at an upper bound rather than an accurate measure of judgment.

Lu: It’s fascinating because it highlights that there are different ways to look at the same AI output, and these two methods yield very different results when we test them against position bias.

Meng: I see a practical implication here that we need to be very careful about which readout method we use depending on what kind of accuracy we actually need for our systems.

Lalam: This research really suggests that improving how AI outputs are evaluated means recognizing these subtle ways simple extra steps can introduce significant measurement errors, and that’s something we should focus on.

Tom: And this leads us to think about the bigger picture, because if we understand this distortion in how we judge AI decisions, it opens up new avenues for building more trustworthy and reliable AI tools.

Jane: Indeed, because understanding these measurement pitfalls means we can design systems that are built on a foundation of more honest and complete evaluation data.

Lu: The future work suggested by the authors seems to be exploring how this adaptive computation could be integrated into deployment strategies to mitigate this risk while still keeping things efficient.

Meng: I'm interested in seeing how these adaptive methods actually translate into real-world performance gains, because a theoretical fix isn't as good as something that runs smoothly on a live system.

Lalam: And ultimately, it gives us a clearer path toward building AI where we can be fully transparent about the judgment process itself and improve the culture around trust in these models.

Gnaneswar Villuri Hashmath Shaik Alex Doboli

Department of Electrical and Computer Engineering, Stony Brook University

cs.CL, cs.AI

Submitted: 2026-09-04

Updated: 2026-09-04

Code: https://github.com/EleutherAI/lm-evaluation-harness

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 87/100

The gist: Reading an LLM judge’s verdict from its first generated token reveals that this cheap readout distorts position bias in one direction, meaning figures obtained this way behave as upper bounds

Key concepts

Position Bias
This refers to a tendency in LLM judging where the model's position in a pair of responses influences its verdict. It is measured by checking how often the verdict flips when the two response orders are swapped. The paper shows that reading only the first token distorts this bias significantly.
Forced Read
This technique involves reading the verdict from the logits of only the very first generated token, without generating any subsequent text. It is a cheap way to get a verdict but is unreliable because judges do not always start with a definitive answer, leading to errors in judgment.
Protocol Divergence
The paper compares two ways of getting a verdict: one involves parsing the text after generation, and the other involves reading the first token's logits. These methods yield different results. The divergence is most noticeable when judges don't start with a clear verdict token.
Compliance Rate
This measures how often a judge actually commits to a verdict for a given pair of responses. It costs one extra computation per pair and varies widely across different conditions, from 0.506 to 1.000. This rate is crucial because the two readout methods yield different quantities when compliance is incomplete.

Terminology

Summary

Reading an LLM judge’s verdict from its first generated token reveals that this cheap readout distorts position bias in one direction, meaning figures obtained this way behave as upper bounds rather than accurate measures of judgment.

The gist

Reading an LLM judge’s verdict from the logits of its first generated token is cheap, requires no generation, and is exactly what constrained decoding and likelihoodscoring evaluation harnesses produce. We show that this readout distorts position bias in one direction: it overstates it in every condition we test, so figures obtained this way behave as upper bounds. The mechanism is that judges do not always lead with a verdict token, on 12% to 49% of pairs for three Qwen3 judges and under 3% for Llama-3.1-8B and Phi-3.5-mini, and forcing a read on those pairs returns whichever response was shown first rather than a judgment. Pooled over the 924 pairs where a judge did not commit, the forced read flips on 89.7% of them when the responses are swapped, against 47.5% read after generation (paired difference +0.422, 95% CI [+0.365, +0.467]). The distortion is specific to what is measured: it moves position bias by 42 points while moving judge accuracy by under one point in seven of ten conditions, so it misleads whoever audits a judge rather than whoever uses one.

Position Bias Measurement and Protocols

Position bias is measured by presenting each pair of responses in both orders and counting how often the verdict flips, which requires choosing a protocol. The first protocol generates a response and parses the verdict from the text, which is what judge audits do. The second protocol reads the verdict from the logits of the first generated token; it is far cheaper because it requires no generation at all, it is what restricted-likelihood evaluation harnesses compute, and it is what constrained decoding produces by construction, since token masking guarantees a legal verdict at the first position.

How Readouts Diverge

The two protocols in common use are not the same operation. Existing practice handles a non-verdict first token in three ways: 1) the forced read, 2) restricted-likelihood harnesses (which score only candidate answers against each other and never ask what the model would have said), and 3) generating and, on a parse failure, retrying or letting the judge continue. The paper focuses on the forced read, which takes the larger of two verdict-token logits at the first generated position. This property is what this paper is about.

Key Findings on Distortion

The divergence between protocols is most pronounced when a judge did not lead with a verdict token. Pooled over 924 such pairs where a judge did not commit, the forced read flips on 89.7% of them when the responses are swapped, against 47.5% read after generation (paired difference +0.422). This distortion is specific to what is measured: it moves position bias by 42 points while moving judge accuracy by under one point in seven of ten conditions, so it misleads whoever audits a judge rather than whoever uses one.

Recommendations for Disclosure

The characterisation suggests a shortcut: read the first token, and generate only for the orders where the judge did not lead with a verdict. This adaptive computation is standard but those cascades escalate to a stronger model under a calibrated threshold, whereas escalation between two readouts of the same model needs no calibration set. The saving is large and fidelity is high but not perfect. The paper recommends reporting three things in judge deployment disclosure: 1) the compliance rate, which costs one forward pass per pair and varies from 0.506 to 1.000 across conditions; 2) which readout produced the verdict, because figures obtained by constrained decoding and generate-and-parse are not the same quantity when compliance is incomplete; and 3) treating the answer format as an experimental variable, noting that changing two characters moves compliance by 17 points here. The first token is a convenient place to read a judge’s verdict and an unreliable one, as it is not a verdict at all on up to 49% of pairs.

Limitations

The divergence is measurable in only one of the three families tested, as compliance is so high for the other two that Phi-3.5-mini contributes only 18 non-compliant pairs and Llama-3.1-8B one, too few to analyze for a pooled comparison. The study also tests only open-weight judges at 8B and below because frontier API judges do not expose logits. Furthermore, the generated readout uses a heuristic parser and a capped budget, making it a lower bound on how often a verdict is eventually reached.

Improvements for AI systems

Here are the specific improvements that can be made to AI systems, based on the findings of this research:

  1. Use a hybrid judging protocol for LLM judges, specifically leveraging constrained decoding or restricted-likelihood evaluation harnesses, instead of relying solely on reading the logits of the first generated token. This addresses the distortion where forced reads overstate position bias by moving it in one direction (overstating it).

  2. Implement a conditional generation strategy for auditing: Only generate responses for those specific orders where the judge did not lead with a verdict token, and use this filtered set to perform audit comparisons. This significantly reduces computational cost while maintaining high fidelity in error measurement compared to generating full responses or relying solely on the forced read readout.

  3. Report Compliance Rate as a mandatory metric in judge deployment disclosures. This rate (the fraction of pairs where the judge would have led with a verdict token in both orders) is a model-specific property that cannot be inherited from other studies and should be reported alongside position bias figures to provide an accurate picture of the judge's reliability.

  4. Distinguish between two distinct measurement quantities:

Choose one readout protocol (forced read vs. generated readout) based on the specific audit goal, recognizing that they measure different things (position bias vs. judge accuracy).

  1. Treat the answer format itself as an experimental variable during evaluation. Since changing just two characters can move compliance by 17 points, AI system evaluations should explicitly test and report which verdict tokens were used to understand how output formatting influences judge behavior, rather than treating the output format as a fixed input.

Abstract

Reading an LLM judge's verdict from the logits of its first generated token is cheap, requires no generation, and is exactly what constrained decoding and likelihood-scoring evaluation harnesses produce. We show that this readout distorts position bias in one direction: it overstates it in every condition we test, so figures obtained this way behave as upper bounds. The mechanism is that judges do not always lead with a verdict token, on 12% to 49% of pairs for three Qwen3 judges and under 3% for Llama-3.1-8B and Phi-3.5-mini, and forcing a read on those pairs returns whichever response was shown first rather than a judgment. Pooled over the 924 pairs where a judge did not commit, the forced read flips on 89.7% of them when the responses are swapped, against 47.5% read after generation (paired difference +0.422, 95% CI [+0.365, +0.467]). The distortion is specific to what is measured: it moves position bias by 42 points while moving judge accuracy by under one point in seven of ten conditions, so it misleads whoever audits a judge rather than whoever uses one. A second, smaller failure occurs even when the judge does lead with a verdict token, since it sometimes opens with one letter and reasons its way to the other, on 0 to 5.5% of pairs at a rate uncorrelated with compliance. We recommend reporting the rate at which a judge leads with a verdict token, which costs one forward pass and no labels, alongside any position-bias figure.

Sources

Related papers