Where Do Apparent LLM Clinical Triage Failures Arise? Localizing the Multiple-Choice Format Effect
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Where Do Apparent LLM Clinical Triage Failures Arise? Localizing the Multiple-Choice Format Effect".
Jane: Patient-voiced clinical-triage benchmarks report high under-triage rates for consumer LLMs for constrained multiple-choice output, yet the same cases score differently with free-text.
Tom: First, who's behind it and why it matters.
Paper summary: Tom: Now that we’ve touched on what they found, I want to go into a bit more detail about the core argument of this paper, "Where Do Apparent LLM Clinical Triage Failures Arise? Localizing the Multiple-Choice Format Effect." Essentially, the researchers are tackling the issue where consumer LLMs show high under-triage rates when forced into constrained multiple-choice outputs.
Jane: That's a tough spot for any clinical application, Tom, because those apparent failures can look like a total breakdown in understanding patient risk. The paper’s thesis is quite specific: they are asking whether the output format itself is what changes how the model represents clinical information or if it’s just changing the mapping from that representation to a final answer choice.
Lu: They set up two hypotheses to test this: an encoding-shift hypothesis, which suggests format prevents knowledge access, and an output-mapping hypothesis, which suggests the format just alters how a pre-existing representation is mapped to a result.
Meng: So they are trying to separate whether the model can't "see" the medical facts under multiple choice or if it just doesn't know how to translate those facts into one of the four options correctly.
Lalam: It’s like asking if a student forgets a fact because they can’t access it, or if they get stuck because the test only allows answers in a specific shape.
Tom: Exactly! And their findings strongly support the output-mapping hypothesis, showing that medical content is preserved across multiple-choice and free-text formats as it peaks on the clinical narrative in both cases.
Jane: That’s a huge piece of evidence; it confirms that the underlying clinical reasoning isn't degraded when we look at the raw text, which is really reassuring for many people.
Lu: The mechanistic localization they found using three independent methods—autoencoder verbalization, decision-token logit attribution, and top-feature characterization—all pointed to a singular conclusion about where the shift occurs.
Meng: They found that medical features fire on the shared clinical narrative in both formats but then go silent at the decision token when multiple choice is involved.
Lalam: That means we can isolate this specific layer or token as the critical juncture for format-induced attention shifting, which is a very precise finding.
Tom: And they also found that scaffold features dominate the logits at that decision point, while medical features are effectively ignored there, even when looking at models like Gemma three 4B IT and Qwen3-8B <ref:2605.29889#pg0>.
Jane: That’s the mechanism: the model is shifting its focus from the actual patient symptoms to structural or contextual features right before it commits to an answer choice.
Lu: They also noted that the gap between formats isn't due to deferral, which is a label-space concern, but rather miscalibration, which they found dominates across every model tested.
Meng: So the implication for engineering is that we need to focus on improving how the model maps its internal state at that final decision point, not just on input quality.
Lalam: This suggests that even if our models have all the clinical knowledge stored, the way we prompt or structure the final output forces a representation change that needs careful calibration.
Conclusion: Tom: So, wrapping up this discussion on "Where Do Apparent LLM Clinical Triage Failures Arise? Localizing the Multiple-Choice Format Effect," the authors are showing us that those seemingly disastrous under-triage rates we see in consumer LLMs for multiple-choice triage actually stem from a format effect.
Jane: That’s the central message, Tom: it’s not that the AI is suddenly less smart clinically; it's that when you constrain its output to a specific format like multiple choice, its internal focus shifts away from the core medical data toward other features.
Lu: The authors use sparse autoencoder features to provide concrete evidence, showing that while medical information remains present in the narrative, it doesn't activate at the decision point under those constrained conditions.
Meng: From a practical viewpoint, this means we don't have to overhaul our entire clinical reasoning engine just because we change how we ask for an answer; instead, we target the mapping layer where that feature silencing happens.
Lalam: It’s about building systems where the most critical medical signals maintain their saliency right up to the moment of decision, no matter what format you impose on them later.
Tom: The authors conclude that this effect is localized to output-mapping, meaning medical content is preserved upstream in the narrative, but the format dictates which feature set gets activated at that final decision point.
Jane: So, in simple terms for our listeners, it means if we want better triage accuracy, we shouldn't just focus on making the model more knowledgeable; we need to ensure its internal mechanism prioritizes clinical facts consistently across different output structures.
Lu: The implication for the future is that researchers will be able to design architectures that are explicitly robust against this specific type of feature silencing at decision points, leading to more reliable performance regardless of whether they use a list or a paragraph format.
Meng: I think this points toward designing more flexible interfaces where the model can choose its representation based on the task, rather than being rigidly locked into one structure that causes these internal shifts.
Lalam: For culture, it’s about establishing a standard where critical medical features are always prioritized at the most sensitive stage of AI output generation, ensuring safety regardless of presentation.
David Fraile Navarro, Berardino Como, Jialei Sheng, Soundariya Ananthan, Shlomo Berkovsky
Macquarie University · Politecnico di Bari
cs.CL, cs.AI
Submitted: 2026-05-28
Updated: 2026-10-02
Code: https://github.com/dafraile/SAE_mad
Importance score: 83/100
The gist: Patient-voiced clinical-triage benchmarks report high under-triage rates for consumer LLMs for constrained multiple-choice output, yet the same cases score differently with free-text.
Key concepts
- Output-Mapping Hypothesis
- This hypothesis suggests that the difference in performance between formats arises because the output format dictates how a pre-existing clinical representation is mapped onto a specific answer. The underlying medical knowledge remains intact and preserved throughout the process, but the final decision token highlights different features based on whether it's structured or free text.
- Decision Token Logit Attribution
- This method was used to pinpoint exactly which features influence the model's final choice at the decision token. The results showed that medical features—the actual clinical content—become silent at this point, while scaffold features, which relate to the structure of the prompt or options, become dominant in determining the final letter chosen.
- Single-Acuity-Step Miscalibration
- This explains why free-text and multiple-choice scores differ. The model isn't failing to access knowledge entirely; instead, it systematically picks an answer that is only one step away from the correct one in terms of acuity, rather than completely missing the correct category.
Terminology
Summary
Patient-voiced clinical-triage benchmarks report high under-triage rates for consumer LLMs for constrained multiple-choice output, yet the same cases score differently with free-text. The core finding is that apparent multiple-choice penalty in LLM clinical triage originates from an output format effect rather than a degradation of underlying clinical reasoning.
The Gist
The apparent multiple-choice penalty in LLM clinical triage originates from a representational shift at the decision token, where medical features go silent and scaffold features dominate, while medical content is preserved upstream.
Investigating the Format Effect
The researchers investigated whether output format changes the model’s clinical representation or only the mapping from a preserved representation to an answer. They tested two hypotheses: an encoding-shift hypothesis
(failure to access knowledge under multiple-choice format) and an output-mapping hypothesis
(format changes how the model maps a pre-formed representation). The study found that medical content is preserved across formats, as the identified medical features peak on the clinical narrative in both multiple-choice and free-text cases.
Mechanistic Localization of Failure
Three independent methods—natural-language autoencoder verbalization, decision-token logit attribution, and top-feature characterization—agreed that scaffold and format features drive decision logits, but not medical features. Specifically:
-
Medical features fire on the shared clinical narrative under both formats and peak inside the vignette in 98–100% of (case, feature) combinations across three instruction-tuned LLMs (Gemma 3 4B IT, Gemma 3 12B IT, and Qwen3-8B).
-
At the multiple-choice decision token, medical features go silent and scaffold features take over. This was corroborated by SAE logit attribution and top-20 active feature characterization analyses.
-
The gap between formats is decomposed into miscalibration (dominating at every model) and deferral (a label-space concern not contributing to the measured gap).
Behavioral Analysis and Robustness Checks
The study analyzed behavioral differences across four conditions: Structured + Multiple-Choice (SL), Natural + Multiple-Choice (NL), Structured + Free-Text (SF), and Natural + Free-Text (NF).
-
The multiple-choice penalty inverts under both structured and natural-language input, ruling out positional bias via option order shuffles.
-
The gap between multiple-choice and free-text formats is dominated by
single-acuity-step miscalibration,
where the model systematically picks an adjacent acuity letter to the gold answer rather than failing to access knowledge. -
A linear-probe framework was deployed to predict which cases will flip correctness between formats from source-format hidden states alone, showing that NL-sourced last-token embeddings produce the most consistent predictive signal across target directions at late layers.
Invariance and Feature Characterization
The study employed a medical vs. non-medical
contrastive test using SAE features to determine if medical features show greater format invariance than random features.
-
Medical features are significantly more format-invariant than magnitude-matched random features under both the natural (NL-NF) and structured (SL-SF) input pairs, with the difference being robust across feature set sizes (K=3 to K=20).
-
The direction of the residual difference in SAE basis projects onto non-medical scaffold features, indicating that format direction lives primarily in non-medical features.
-
Decision-token analysis confirmed that medical features have zero activation at the decision token under both formats, while scaffold features drive the letter logits.
Conclusion and Clinical Implications
The four converging observations support an output-mapping reading: medical content is preserved on the clinical narrative, but format dictates which feature set (medical vs. scaffold) is active at the decision point. The gap is driven by single-acuity-step miscalibration, not deferral, suggesting the benchmark's 4-letter space cannot express conditional triage boundaries but does not reflect a failure of underlying clinical reasoning or safety in free-text outputs. The paper concludes that the format effect is localized to the output stage mapping rather than degraded clinical reasoning as the source of apparent failure rates.
Key Findings Summary
(Note: Specific quantitative results are summarized by referencing key findings from Tables 1, 4, and 9.)
-
Medical features peak on the clinical narrative in both multiple-choice and free-text formats.
-
At the decision token, medical features go silent while scaffold features dominate the logits.
-
The format effect is localized to scaffold features at the decision token while medical content remains intact upstream (output-mapping hypothesis).
-
The NL-NF accuracy gap is dominated by single-acuity-step miscalibration, not deferral.
Improvements for AI systems
As a fastidious researcher, I have analyzed this paper, Internal Representation, Not Clinical Knowledge: Where Apparent LLM Triage Failures Originate.
The core finding is that apparent triage failures in consumer LLMs stem from an output-mapping effect rather than degraded clinical reasoning; specifically, the model's attention shifts from medical content to scaffold (multiple-choice) features at the decision token.
Here are specific, actionable improvements for AI systems derived from this research:
) Specific Improvements and Capabilities of an Improved System:
The system should incorporate a dual-path verification mechanism that distinguishes between clinical content understanding and prompt structure dependency.
Implement a Format-Aware Attention Check
at the decision token level to flag when scaffold features are dominating over medical features during answer selection, allowing for automated fallback or re-prompting if this shift is detected.
Develop an input-style robustness suite that tests model performance across varying input formats (structured vs. natural language) while holding the core clinical content constant, ensuring the system isn't brittle to prompt style changes.
Deploy a Feature Silence Detection
module trained on SAE features to monitor medical feature firing during answer selection; if these features go silent while scaffold features dominate, the system should recognize this as a potential representation shift and adjust its output strategy accordingly.
Integrate a Positional Bias Mitigation Layer
that actively shuffles answer options (or re-orders them) in cases where the model shows positional bias, thereby forcing it to rely on content priors rather than the rigid layout of the multiple-choice scaffold.
Utilize a Constraint-First Prompting Strategy
during inference for high-stakes tasks, moving formatting instructions to the beginning of the prompt to isolate and minimize representational shifts caused by positional artifacts in later prompt elements.
Establish a Clinical Content Anchor
mechanism that verifies medical feature activations are consistently anchored within the shared clinical narrative tokens, ensuring that when a model makes a decision, it is grounded in the actual patient symptoms rather than superficial formatting cues.
) What the Improved AI System Can Do:
The improved system will transition from being a simple pattern matcher to an Internal Representation Auditor.
It can perform the following specific functions:
Format-Aware Triage
: When presented with a triage question in multiple-choice format, the system will internally monitor its attention weights at the decision point. If it detects that medical features are ceasing to fire and scaffold features are taking over, it will flag this as a Representation Shift Alert
and trigger a safety protocol (e.g., requesting clarification or defaulting to a conservative assessment).
Robustness-Driven Triage
: The system will be explicitly trained not just on correct answers, but on the invariance of its triage decisions across input styles (structured vs. natural language). This makes it far more reliable when deployed in real-world settings where user queries arrive in varied formats.
Bias-Resistant Response Generation
: By incorporating positional bias mitigation, the system will be less likely to make incorrect choices based on the arbitrary ordering of options, leading to more stable and content-driven triage decisions across different question sets.
Grounding Verification
: The Clinical Content Anchor ensures that any output generated is directly traceable back to the medical facts presented in the patient vignette, effectively preventing hallucinated
or format-driven answers that lack clinical substance.
Sources
- Interpretability without actionability: mechanistic methods cannot correct language model errors despite near-perfect internal representations
- Sparse Autoencoders Find Highly Interpretable Features in Language Models
- Evaluation format, not model capability, drives triage failure in the assessment of consumer health AI
- Scaling and evaluating sparse autoencoders
- The Curious Case of Neural Text Degeneration
- Language Models (Mostly) Know What They Know
- Towards Principled Evaluations of Sparse Autoencoders for Interpretability and Control
- k-Sparse Autoencoders
- Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models
- Steering Language Models With Activation Engineering
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering