Where Do Apparent LLM Clinical Triage Failures Arise? Localizing the Multiple-Choice Format Effect

summary

Video file (mp4)

The gist

Patient-voiced clinical-triage benchmarks report high under-triage rates for consumer LLMs for constrained multiple-choice output, yet the same cases score differently with free-text.

In short

The study investigated why LLMs perform worse in clinical triage when forced to choose from multiple-choice options compared to free-text responses. The core finding is that this penalty isn't due to poor clinical reasoning but an output format effect. Medical content is preserved, but the model shifts its focus at the final decision point from medical features to scaffold features, leading to miscalibration in choosing between adjacent acuity levels.

Key concepts

Output-Mapping Hypothesis
This hypothesis suggests that the difference in performance between formats arises because the output format dictates how a pre-existing clinical representation is mapped onto a specific answer. The underlying medical knowledge remains intact and preserved throughout the process, but the final decision token highlights different features based on whether it's structured or free text.
Decision Token Logit Attribution
This method was used to pinpoint exactly which features influence the model's final choice at the decision token. The results showed that medical features—the actual clinical content—become silent at this point, while scaffold features, which relate to the structure of the prompt or options, become dominant in determining the final letter chosen.
Single-Acuity-Step Miscalibration
This explains why free-text and multiple-choice scores differ. The model isn't failing to access knowledge entirely; instead, it systematically picks an answer that is only one step away from the correct one in terms of acuity, rather than completely missing the correct category.

Terminology used across episodes

This episode discusses

The paper

Where Do Apparent LLM Clinical Triage Failures Arise? Localizing the Multiple-Choice Format Effect · Read on arXiv

David Fraile Navarro, Berardino Como, Jialei Sheng, Soundariya Ananthan, Shlomo Berkovsky

Macquarie University · Politecnico di Bari

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Where Do Apparent LLM Clinical Triage Failures Arise? Localizing the Multiple-Choice Format Effect".

Jane: Patient-voiced clinical-triage benchmarks report high under-triage rates for consumer LLMs for constrained multiple-choice output, yet the same cases score differently with free-text.

Tom: First, who's behind it and why it matters.

Paper summary: Tom: Now that we’ve touched on what they found, I want to go into a bit more detail about the core argument of this paper, "Where Do Apparent LLM Clinical Triage Failures Arise? Localizing the Multiple-Choice Format Effect." Essentially, the researchers are tackling the issue where consumer LLMs show high under-triage rates when forced into constrained multiple-choice outputs.

Jane: That's a tough spot for any clinical application, Tom, because those apparent failures can look like a total breakdown in understanding patient risk. The paper’s thesis is quite specific: they are asking whether the output format itself is what changes how the model represents clinical information or if it’s just changing the mapping from that representation to a final answer choice.

Lu: They set up two hypotheses to test this: an encoding-shift hypothesis, which suggests format prevents knowledge access, and an output-mapping hypothesis, which suggests the format just alters how a pre-existing representation is mapped to a result.

Meng: So they are trying to separate whether the model can't "see" the medical facts under multiple choice or if it just doesn't know how to translate those facts into one of the four options correctly.

Lalam: It’s like asking if a student forgets a fact because they can’t access it, or if they get stuck because the test only allows answers in a specific shape.

Tom: Exactly! And their findings strongly support the output-mapping hypothesis, showing that medical content is preserved across multiple-choice and free-text formats as it peaks on the clinical narrative in both cases.

Jane: That’s a huge piece of evidence; it confirms that the underlying clinical reasoning isn't degraded when we look at the raw text, which is really reassuring for many people.

Lu: The mechanistic localization they found using three independent methods—autoencoder verbalization, decision-token logit attribution, and top-feature characterization—all pointed to a singular conclusion about where the shift occurs.

Meng: They found that medical features fire on the shared clinical narrative in both formats but then go silent at the decision token when multiple choice is involved.

Lalam: That means we can isolate this specific layer or token as the critical juncture for format-induced attention shifting, which is a very precise finding.

Tom: And they also found that scaffold features dominate the logits at that decision point, while medical features are effectively ignored there, even when looking at models like Gemma three 4B IT and Qwen3-8B <ref:2605.29889#pg0>.

Jane: That’s the mechanism: the model is shifting its focus from the actual patient symptoms to structural or contextual features right before it commits to an answer choice.

Lu: They also noted that the gap between formats isn't due to deferral, which is a label-space concern, but rather miscalibration, which they found dominates across every model tested.

Meng: So the implication for engineering is that we need to focus on improving how the model maps its internal state at that final decision point, not just on input quality.

Lalam: This suggests that even if our models have all the clinical knowledge stored, the way we prompt or structure the final output forces a representation change that needs careful calibration.

Conclusion: Tom: So, wrapping up this discussion on "Where Do Apparent LLM Clinical Triage Failures Arise? Localizing the Multiple-Choice Format Effect," the authors are showing us that those seemingly disastrous under-triage rates we see in consumer LLMs for multiple-choice triage actually stem from a format effect.

Jane: That’s the central message, Tom: it’s not that the AI is suddenly less smart clinically; it's that when you constrain its output to a specific format like multiple choice, its internal focus shifts away from the core medical data toward other features.

Lu: The authors use sparse autoencoder features to provide concrete evidence, showing that while medical information remains present in the narrative, it doesn't activate at the decision point under those constrained conditions.

Meng: From a practical viewpoint, this means we don't have to overhaul our entire clinical reasoning engine just because we change how we ask for an answer; instead, we target the mapping layer where that feature silencing happens.

Lalam: It’s about building systems where the most critical medical signals maintain their saliency right up to the moment of decision, no matter what format you impose on them later.

Tom: The authors conclude that this effect is localized to output-mapping, meaning medical content is preserved upstream in the narrative, but the format dictates which feature set gets activated at that final decision point.

Jane: So, in simple terms for our listeners, it means if we want better triage accuracy, we shouldn't just focus on making the model more knowledgeable; we need to ensure its internal mechanism prioritizes clinical facts consistently across different output structures.

Lu: The implication for the future is that researchers will be able to design architectures that are explicitly robust against this specific type of feature silencing at decision points, leading to more reliable performance regardless of whether they use a list or a paragraph format.

Meng: I think this points toward designing more flexible interfaces where the model can choose its representation based on the task, rather than being rigidly locked into one structure that causes these internal shifts.

Lalam: For culture, it’s about establishing a standard where critical medical features are always prioritized at the most sensitive stage of AI output generation, ensuring safety regardless of presentation.

More episodes

← Home