The Knowing-Saying Gap: When Probes See Errors that Confidence Misses

summary

Video file (mp4)

The gist

The paper "The Knowing-Saying Gap: When Probes See Errors that Confidence Misses" investigates a dissociation in large language models: linear probes on internal activations detect corrupted context

In short

The paper 'The Knowing-Saying Gap' examines a fundamental disconnect in AI models where internal mechanisms detect errors but fail to reflect them in output confidence. Researchers found that internal probes identify corruption with high accuracy, yet the model's verbalized certainty remains constant. This gap is observed across multiple model families and challenges assumptions about how to monitor and improve AI reliability.

Key concepts

Knowing-Saying Gap
This refers to the dissociation between what a model internally computes (its hidden layers or 'knowing') and what it actually outputs (the words it says). The paper demonstrates that models can internally flag an error while maintaining high confidence in their final, potentially incorrect, answer.
Linear Probe
A tool used by researchers to measure a model's internal state. These probes detect corruption within the model's activations with near-perfect accuracy (AUROC above 0.98) but are useless at predicting if the final output will be wrong.
Fluency without Fidelity
This term describes the central reliability problem in deployed language models. The models generate fluent, confident responses even when their internal reasoning is flawed or corrupted, lacking the necessary fidelity to reflect that uncertainty.
Interventions
Methods used to correct errors in a model's process. These include Reprompt (adding a self-check message), Replace-prior (regenerating and swapping a previous step), and Branch-and-pick (sampling candidates and using the internal probe to select the least corrupted one).

Terminology used across episodes

This episode discusses

The paper

The Knowing-Saying Gap: When Probes See Errors that Confidence Misses · Read on arXiv

Jyotin Goel, Ipshita Bandyopadhyay, Justin Shenk

Linear probes detect corrupted context in language models with near-perfect accuracy, yet this does not translate into reliable failure prediction. The result is a dissociation with direct implications for deployment monitoring. Across multi-hop arithmetic chains, probes that detect corruption turn out to be uninformative about final answer correctness; models forced into structured confidence formats collapse to two values with indistinguishable error rates; and probe persistence across hops fails to separate correct from incorrect outcomes, refuting our pre-registered "persistence beats peak" hypothesis. This pattern of knowing but not saying generalises across model families including reasoning models. As a real-time monitor, probe-based interventions are sharply model and error-type dependent: branch-and-pick is net-positive across models and uniquely non-breaking on Llama-3.1-8B (4 rescued, 0 broken), while reprompt and replace-prior break correct traces at roughly the rate they rescue wrong ones. Probe-based monitoring is a necessary complement to verbalised confidence, but no single intervention dominates, and the deployable answer is model-aware, error-type-aware routing.

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "The Knowing-Saying Gap: When Probes See Errors that Confidence Misses".

Jane: The paper was written by Jyotin Goel, Ipshita Bandyopadhyay and Justin Shenk from.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome back to the show, everyone. We've got a paper that's been making the rounds, and the title alone grabbed me: "The Knowing-Saying Gap: When Probes See Errors that Confidence Misses." Jane, what's your first read on that?

Jane: Tom, it's one of those titles that sounds like a riddle, but it's actually pointing at something really concrete. It's about the difference between what a model has internally computed and what it actually tells us. You know, the "knowing" part is what's in the hidden layers, and the "saying" part is the words it outputs.

Tom: And the paper is basically saying those two things can completely disconnect. The model can internally flag that something's wrong, but when it speaks, it's just as confident as ever. No hedging, no slowdown, no "wait, let me check that."

Jane: Exactly. And that's a terrifying thought for anyone deploying these models in the real world. If a coding agent is on step one of a five-step task and it makes a mistake, that error is now baked into the context for every subsequent step. The paper calls it a "seed" — a wrong answer that grows into a corrupted chain.

Tom: Right, and the authors set up this really clean experiment to isolate it. They built a dataset of multi-hop arithmetic problems, injected silent errors into the context, and then checked two things: can a probe on the internal activations detect the corruption, and does that detection actually predict whether the final answer will be wrong?

Jane: And here's the kicker. The probe detects the corruption with near-perfect accuracy — we're talking AUROC above zero point nine eight across all five models they tested. But that same probe is completely useless at predicting whether the final answer is wrong. It's literally at chance, like zero point five.

Tom: So the model knows the context is corrupted, but that knowledge doesn't propagate to the output. It's like a pilot whose instruments show a warning light, but the autopilot just keeps flying straight into the mountain.

Jane: That's the "knowing-saying gap" in a nutshell. And the paper shows it's not just one model — it's across Qwen and Llama families, across instruction-tuned and even reasoning models. The reasoning model actually made it worse, which we'll get into.

Tom: I'm already hooked. But before we go deeper, I want to bring in Lu from Tsinghua, because I know this kind of result is going to spark some ideas. Lu, what's your gut reaction?

Lu: Tom, my first thought is that this reframes the entire calibration literature. We've spent years trying to get models to tell us when they're uncertain, but this paper suggests the verbal channel is fundamentally broken for this kind of error. The information is there, linearly decodable, but the generation policy just doesn't consult it.

Jane: And that's the part that's going to keep researchers busy for a while. It's not that the model can't know — it's that the knowing doesn't make it to the saying. We'll dig into what that means for interventions next.

Summary: Tom: So we've established the core finding — probes see the error, confidence doesn't. But what does the paper actually do with that? Jane, walk us through the summary.

Jane: Well, Tom, they run a full pipeline. They collect activations, train a linear probe, measure the detection-versus-failure dissociation, and then they test three different interventions to see if the probe signal can actually be used to fix things in real time.

Tom: And the interventions are where it gets really interesting, because they're not all created equal. We've got reprompt, replace-prior, and branch-and-pick. Can you break those down for our listeners?

Jane: Sure. Reprompt is the cheapest — you just add a message saying "recheck your work" and hope the model catches its own mistake. Replace-prior is more aggressive — you regenerate the previous hop from scratch and swap it into the context. And branch-and-pick is the fancy one — you sample multiple candidate responses and use the probe to pick the one that looks least corrupted internally.

Tom: And the results are honestly a bit sobering. Reprompt does essentially nothing. Replace-prior breaks correct traces at almost the same rate it rescues wrong ones. Only branch-and-pick is consistently net-positive, and even then it's not a silver bullet.

Lu: What struck me, Tom, is that the probe fires on ninety-six to one hundred percent of error traces, but replace-prior's firing drops to thirty-two to thirty-nine percent when you actually use it. That means the probe is detecting the corruption, but removing the corrupted value doesn't fix the answer. The damage is already done downstream.

Jane: Right, and that's a really important nuance. The probe is decodable, but it's not necessarily causal. The paper is careful to say they're showing decodability, not causation. The error is encoded in the activations, but that doesn't mean the generation policy is using that encoding to produce the wrong answer.

Tom: And then there's the thinking-mode result, which I think is going to surprise a lot of people. They ran the same model with chain-of-thought reasoning enabled, and the probe still detects the error just as well — but the accuracy on error traces collapses from about twenty-eight percent down to one point two percent.

Jane: So the reasoning model is generating five times more tokens, thinking harder, and yet it's worse at recovering from the injected error. And it's still internally aware of the corruption. That's the knowing-saying gap on steroids.

Meng: As the engineer in the room, I have to ask — what does this mean for actually deploying these monitors? If branch-and-pick is the only safe option, and it costs roughly ten times more inference compute, that's a real budget question.

Tom: That's exactly where the paper lands. They're not claiming a magic fix. They're saying probe-based monitoring is a necessary complement to verbalized confidence, but the deployable answer is model-aware, error-type-aware routing. No single intervention dominates.

Jane: And we're going to get into those per-error-type results in a bit, because that's where the nuance really lives. But first, let's talk about what's on the first page, because the framing there is crucial.

Improvements: Tom: Alright, so we've covered the headline findings. But what does the paper suggest we actually do about this? Jane, what's the improvement story here?

Jane: The paper's core suggestion is that we need to stop relying on verbalized confidence as our primary monitoring channel. Instead, we should be looking at internal state — specifically, linear probes on the residual stream — as a necessary complement. The authors are pretty direct about this: trusting the verbal channel is trusting a channel they've shown to be uninformative.

Tom: And they back that up with a head-to-head comparison. They tested six different surface signals — things like entropy, log-probability, vocabulary gap — and none of them got above zero point seven zero AUROC for detection. Meanwhile, the probe is sitting at zero point nine eight or higher. That's a massive gap.

Lu: What I find compelling, Tom, is the robustness work they did. They ran grouped cross-validation, holding out entire base problems, and the detection barely moved — within zero point zero zero three AUROC. Then they did leave-one-error-type-out, training on four error types and testing on a fifth, and it still transferred above zero point nine one in every single cell.

Jane: That's the part that convinces me this isn't a surface artifact. The probe isn't memorizing specific error tokens. It's encoding an abstract "something is wrong here" state that generalizes across error types it's never seen.

Tom: So the improvement is really about routing. The paper shows that different interventions work better for different error types. Replace-prior is great for wrong unit errors — it rescues thirty-seven point five percent — but it's net-negative on wrong percentage base. Branch-and-pick is the only strictly non-breaking policy on off by one errors.

Meng: But hold on — that per-type analysis has only nine to sixteen examples per slice. That's thin. Can we really build a router on that?

Jane: You're right to flag that, Meng. The paper is very honest about it. They call routing a "direction" rather than a deployable result. You'd need an error-type classifier first, and they don't have one yet. But the heterogeneity is real, and it's a strong motivation for that future work.

Tom: And there's another improvement angle I want to highlight — the pre-registered hypothesis that got refuted. They thought persistence of probe activation across hops would predict failure. It doesn't. Peak doesn't either. Neither separates correct from incorrect outcomes. That's a null result that saves other researchers from chasing the same dead end.

Lu: I think the biggest improvement suggestion is the one about per-hop localization. Right now the probe says "a prior is corrupted." The next step is to make it say "hop three is corrupted." That would make replace-prior much more targeted and effective.

Jane: And that's the hook for the next segment, because the first page of the paper lays out the real-world stakes that make all of this urgent.

First Page: Tom: So we've talked about the methods and the improvements. But the first page of "The Knowing-Saying Gap" really sets the stage for why this matters. Jane, what stood out to you?

Jane: The opening is about deployed systems. They talk about coding agents like SWE-agent and Devin, and how a misdiagnosed root cause at turn one constrains every subsequent patch. The paper makes the point that even the strongest agents fail on the majority of SWE-bench tasks, and the bottleneck isn't capability — it's cascading context corruption.

Tom: That's a bold claim, but it tracks. If the model makes a wrong assumption early on, every later step is built on sand. And the model doesn't hedge, doesn't slow down, doesn't show any sign that something has gone wrong. It just keeps going in the same confident register.

Jane: And the paper calls that "fluency without fidelity." It's the central reliability problem of deployed language models. And they argue it's going to get worse as models get longer context windows, tool access, and persistent memory. The gap between what's computed and what's said becomes larger and more consequential.

Lu: What I love about the first page is how they frame the natural response. You'd think, "just ask the model." Elicited confidence, chain-of-thought self-critique, structured verification — they all share the assumption that internal state is accessible through verbal outputs. And this paper shows that assumption fails in an informative way.

Tom: "Fails in an informative way" — that's a great phrase. The probe sees the error, but the saying channel is broken. And that's not just a math problem. The paper mentions medical decision support, long-horizon planning, multi-agent coordination. Anywhere a model's output feeds back into its own future context.

Meng: From my side, the first page also makes a practical point about monitoring. A model that stutters on uncertain ground would be easy to watch. But a model that produces wrong answers in the same confident register as correct ones requires a fundamentally different approach. That's the gap this paper is trying to fill.

Jane: And it's a gap that's already visible in production. The authors aren't speculating about future risks — they're documenting a current failure mode. The knowing-saying gap isn't hypothetical. It's happening right now in deployed systems.

Tom: So the first page sets up the problem as urgent and real. And the rest of the paper shows that the solution isn't simple — it's going to require looking inside the model, not just listening to what it says.

Jane: And that's where we're headed next — pulling together what this means for the future of reliable AI deployment.

Conclusion: Tom: Alright, let's wrap this up. We've spent the show on "The Knowing-Saying Gap: When Probes See Errors that Confidence Misses," and I think we've got a clear picture now. Jane, give us the final summary.

Jane: The paper shows a four-part dissociation. First, a linear probe detects injected errors with near-perfect accuracy — above zero point nine eight AUROC across all five models. Second, that same probe is at chance when predicting whether the error reaches the final answer. Third, verbalized confidence collapses to a binary signal with indistinguishable error rates. And fourth, chain-of-thought reasoning widens the gap rather than closing it.

Tom: And the interventions — branch-and-pick is the only consistently net-positive policy, and even that is conditional on error type and model. There's no universal fix.

Lu: What I'll take away is the refutation of the persistence hypothesis. The idea that a longer-lasting probe signal would predict failure turned out to be wrong. That's a useful negative result for the field — it saves us from chasing a plausible-sounding idea that doesn't hold up.

Meng: And from the engineering side, the message is clear: probe-based monitoring is necessary, but it's not sufficient. You need model-aware, error-type-aware routing. And you need to be honest about the cost — branch-and-pick is roughly ten times more expensive in inference compute.

Jane: The paper's final word is that probe-based monitoring is a necessary complement to verbalized confidence, not a replacement. And the deployable answer is routing. That's the direction, even if the router itself doesn't exist yet.

Tom: So we're saying goodbye to "The Knowing-Saying Gap." It's a paper that gives us a new way to think about reliability — not by asking the model to tell us, but by looking at what it actually computed.

Jane: And that's a shift that could change how we monitor every agentic system out there. Thanks for listening, everyone. We'll be back with the next paper soon.

Tom: Until then, keep your probes calibrated and your confidence checks honest. See you next time.

More episodes

← Home