How LLMs Are Persuaded: A Few Attention Heads, Rerouted
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "How LLMs Are Persuaded: A Few Attention Heads, Rerouted".
Jane: A small set of mid-layer attention heads almost entirely determines a language model's answer, and persuasion works by redirecting attention through a rank-one evidence-routing feature.
Tom: First, who's behind it and why it matters.
Title and authors: Tom: So, let's talk about the title and who wrote this paper, "How LLMs Are Persuaded: A Few Attention Heads, Rerouted." It immediately tells us that we are looking at a very targeted investigation into an AI safety concern.
Jane: The authors include Xiangkun Sun from Northeastern University Lingkai Kong, Aoqi Zhang from Tsinghua University Liang Zeng, and Tonghan Wang from Tsinghua University. It shows this work is coming from a strong group focused on the architecture itself.
Lu: I think the combination of researchers here suggests they have deep expertise in both the underlying attention mechanisms and how those mechanisms translate into observable output behaviors under manipulation.
Meng: The title itself points toward a mechanistic understanding, which is crucial because when we deal with model vulnerabilities, we need to know exactly where the fault lies structurally.
Lalam: It’s important to remember that this paper is trying to provide a fundamental explanation for why LLMs can abandon factual knowledge under certain conditions, which moves us past just observing errors and into understanding the cause.
Tom: Precisely; it’s moving from symptom-based fixes to a structural understanding of persuasion, which is where the real progress in safety lies.
Jane: It sets up a really clear roadmap for researchers: identify these specific heads, isolate that routing feature, and then trace back how input keywords construct it.
Lu: The implication here is that we can start thinking about defenses not as broad guardrails, but as surgical interventions targeting these specific attention pathways.
Meng: That level of specificity changes the design process significantly; we can move from general robustness to targeted circuit hardening, which is a huge step for practical deployment.
Lalam: For me, this suggests that the future of AI safety involves mapping out these internal causal chains so we can build defenses that are grounded in the actual math and structure of the model.
The paper's summary: Tom: The core summary of "How LLMs Are Persuaded: A Few Attention Heads, Rerouted" is really about showing that persuasion isn't some kind of general decay in knowledge but a switch governed by a very small group of decision heads.
Jane: They illustrate this by describing how these decision heads encode the four possible answers as vertices of a tetrahedron in a low-dimensional subspace, which helps visualize the geometric space where the model operates.
Lu: That geometric visualization is powerful because it shows that persuasion causes these representations to jump between vertices—from the correct answer vertex to a persuasion target vertex when it succeeds.
Meng: So, they’re establishing that decision heads don't really reason over the evidence; instead, they are simply copying whichever option token their attention happens to select based on an external steering feature.
Lalam: That explains why the model doesn't gradually erode its confidence; it switches answers at rates between twenty-nine and sixty-two percent when presented with persuasive but incorrect context, which is a very specific quantitative finding.
Tom: That quantification is what really hooks me; it moves the discussion away from vague susceptibility to measurable switching rates based on input quality.
Jane: And they emphasize that this override mechanism isn't a fixed property of the model itself, but something constructed dynamically from the input context being fed into it at inference time.
Lu: This dynamic construction is what makes it so hard to defend against static countermeasures because the attack surface is constantly shifting with every prompt.
Meng: If we can’t retrain for every possible persuasion technique, but instead target that rank-one evidence-routing feature, that opens up a whole new category of inference-time interventions.
The paper's improvements: Tom: The paper suggests several ways to build on this discovery, focusing heavily on moving from just understanding the mechanism to actually intervening in it.
Jane: They propose building internal "persuasion monitor" circuits that specifically look for persuasive keywords being read by those identified mid-layer attention heads, like the ones in layers eight through twelve.
Lu: This monitor would trigger a corrective action before an answer is even committed, which I think is a very proactive approach compared to waiting for an error to manifest in the final output.
Meng: From a practical standpoint, implementing this monitoring circuit would require identifying those specific heads first, but if we can do that, it’s much more efficient than broad system checks.
Lalam: They also suggest developing intervention modules that directly target that rank-one option-routing feature derived from the QK circuit to steer the model's choice away from the persuasion target.
Tom: That sounds like a direct attack on the mechanism; modifying or suppressing that specific feature should be much more effective at neutralizing persuasion than general prompt engineering.
Jane: Additionally, they look at integrating a low-dimensional projection layer that maps attention head outputs onto that learned tetrahedral geometry to provide a real-time "persuasion risk" score based on geometric deviation.
Lu: That risk score based on geometric deviation is brilliant because it gives us a continuous metric of how far the model's internal state is drifting from its expected correct path, even when it hasn't made a final error yet.
Meng: I see how that could work in real-time; if the geometric projection shows significant drift towards a persuasion target vertex, we can inject an intervention during inference to pull it back.
Conclusion: Tom: So, to wrap up our discussion on "How LLMs Are Persuaded: A Few Attention Heads, Rerouted," the main point is that persuasion is a discrete jump driven by input-constructed routing features from specific decision heads.
Jane: It’s a compact causal circuit involving shallowness heads reading keywords and writing a feature that redirects attention from the correct option to the persuasion target vertex in their low-dimensional geometric space.
Lu: The implication for us is that we can start building defenses that are mechanistically grounded, not just based on trial and error or symptom patching, by targeting those specific features.
Meng: For engineering teams, this means focusing our efforts on identifying and intervening at the level of the rank-one evidence-routing feature controlled by the QK circuit, which seems like a very precise engineering goal.
Lalam: I see this as a huge cultural shift; it suggests that future AI safety research needs to be deeply integrated with mechanistic interpretability so we can build systems that are inherently resilient to these specific redirection circuits.
Tom: Exactly, it shifts the focus from just trying to make models *behave* safely to understanding precisely *how* they behave when they are steered.
Jane: It’s a really important piece of work because it provides a way for us to move towards targeted runtime monitors and mechanistically grounded defenses against these kinds of manipulations.
Lu: We’ve seen this mechanism appear across different open-source model families and even in scenarios like Generative Engine Optimization, which shows its broad applicability.
Meng: I agree, the transferability across different settings is what makes this finding so robust; it's not just a quirk of one specific training run.
Lalam: And while the paper focuses on controlled multiple-choice settings, we have to remember that the authors themselves noted a limitation: they haven't fully evaluated defenses built upon this circuit yet.
Tom: That’s fair; characterizing the mechanism is step one, but building effective defenses based on those findings is clearly the next major challenge for the community.
Jane: So we’ve looked at how this paper dissects persuasion, and it points us toward a much more precise path for building resilient AI systems.
Northeastern University · Harvard University · Tsinghua University
cs.AI
Submitted: 2026-05-10
Updated: 2026-09-28
Comments: 62 pages, 33 figures, including references and appendices
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 92/100
The gist: A small set of mid-layer attention heads almost entirely determines a language model's answer, and persuasion works by redirecting attention through a rank-one evidence-routing feature.
Key concepts
- Decision Heads
- A small set of mid-layer attention heads that causally determine the model's final answer. They encode the four possible options as vertices in a low-dimensional geometric space, acting primarily as routers rather than deep reasoners.
- Option-Routing Feature
- A one-dimensional feature that controls which option token a decision head selects. This feature is constructed on the fly from persuasive keywords in the input and is the mechanism through which persuasion steers the model's choice.
- Discrete Latent Jump
- The specific effect of successful persuasion, where the model representation switches instantly from being at its correct answer vertex to jumping directly to a predetermined 'persuasion-target vertex' within a geometric space defined by decision heads.
Terminology
Summary
A small set of mid-layer attention heads almost entirely determines a language model's answer, and persuasion works by redirecting attention through a rank-one evidence-routing feature. This mechanism reveals that persuasion is not a diffuse corruption of reasoning but rather a narrow, monitorable circuit involving the rerouting of specific attention heads by input keywords.
The core finding
Persuasion causes a discrete latent jump from the correct-answer vertex to the persuasion-target vertex
within a low-dimensional geometric space defined by decision heads. The model does not gradually erode confidence; instead, it switches answers at rates of 29–62% when presented with persuasive but factually incorrect context. This override is not a fixed property of the model but is constructed on the fly from the input.
Decision Heads and Geometry
The study identifies a sparse set of mid-layer attention heads—decision heads—that causally determine the model’s answer.
These heads encode the four options as vertices of a tetrahedron in a low-dimensional subspace.
The analysis shows that persuasion causes these representations to jump between these vertices. Specifically, when persuasion succeeds, the representation jumps from the original vertex to the persuasion target vertex,
while when it fails, it remains at the correct answer vertex. This geometry is visualized as an approximate regular triangular pyramid.
The Option-Routing Feature
Decision heads do not reason over evidence; they copy whichever option token their attention selects.
The mechanism of persuasion is traced back to a one-dimensional option-routing feature
that controls this selection. This feature is extracted by finding a rank-one approximation of the QK circuit, which governs which key position is selected. Modifying this feature directly steers the model's choice, and removing it blocks persuasion entirely.
The Causal Chain from Input to Output
The paper establishes an end-to-end causal chain from keywords in the input to output.
This chain consists of three verified stages:
-
Persuasive keywords in the input are read by
a band of shallower attention heads
(layers 8–12). -
These shallow heads
write a one-dimensional routing feature uk onto option-token representations.
-
At the decision layer, the QK circuit of the decision head reads this feature, which then redirects attention from the factually correct option to the persuasion target.
Real-World Transfer and Implications
This mechanism is shown to be robust and transferable across different settings. The same mechanism appears across four open-source model families
and in realistic poisoning scenarios such as Generative Engine Optimization (GEO).
Furthermore, the effect of this circuit is preserved when transferring findings from controlled multiple-choice benchmarks (NQ2) to more complex tasks like Geo-Bench, where the selection mechanism determines which candidate source the model favors. This provides a compact causal circuit
that opens the door to building targeted runtime monitors and mechanistically grounded defenses.
Limitations
The study is limited by its focus on controlled multiple-choice settings; whether this exact circuit governs free-form generation remains an open question. The authors also note that while they characterize the mechanism, they do not yet evaluate defenses built upon it. The work is also dual-use, as a mechanistic account could help attackers design more effective persuasive prompts or retrieval-poisoning attacks.
The gist
Persuasion is not a diffuse corruption of reasoning across the network but rather the rerouting of a few mid-layer attention heads by a single one-dimensional feature in the residual stream. Every link in the chain is validated by causal intervention, and this mechanism appears across open-source LLMs and realistic poisoning scenarios such as Generative Engine Optimization, revealing persuasion as a narrow, monitorable circuit.
How it works
-
A few decision heads are identified that causally determine the answer, encoding options as vertices of a tetrahedron in a low-dimensional subspace.
-
Persuasion causes a
discrete jump from the correct-answer vertex to the persuasion-target vertex.
-
Decision heads copy whichever option token their attention selects, meaning they do not compute reasoning over evidence.
-
A rank-one
option-routing feature
governs this selection, which is controlled by a feature constructed on the fly from persuasive keywords in the input. -
Shallow attention heads (layers 8–12) read these persuasive keywords and write the routing feature onto option tokens, completing a verified causal chain from input to output.
Tracing the Mechanism
The researchers traced the override back through three stages: first identifying decision heads as simple routing devices, then isolating a one-dimensional option-routing feature
that governs attention, and finally showing that shallower attention heads construct this feature from persuasive keywords in the input.
This entire chain was validated by causal intervention.
Improvements for AI systems
Here are specific improvements for AI systems based on the findings of this research:
-
Improving Factual Grounding in Retrieval-Augmented Generation (RAG) Systems: Instead of relying solely on prompt engineering or simple fine-tuning, implement a mechanism that dynamically monitors and intervenes in the attention heads responsible for option selection.
-
Developing Causal Monitoring Circuits: Build an internal
persuasion monitor
circuit that detects when persuasive keywords in a retrieved context are read by specific mid-layer attention heads (specifically those identified as being in layers 8–12). This monitor would trigger a corrective action before the final answer is committed. -
Implementing Dynamic Option-Routing Feature Interventions: Develop an intervention module that targets the rank-one option-routing feature (the one derived from the QK circuit) identified as controlling decision heads. By injecting or suppressing this specific feature during inference, the system can prevent choice jumps caused by persuasive input, even if the underlying factual knowledge is correct but susceptible to redirection.
-
Creating Geometric Decision-Space Validators: Integrate a low-dimensional projection layer that maps attention head outputs onto the learned tetrahedral geometry (the decision subspace). This validator would continuously check if an option token's representation is drifting away from its expected vertex, providing a real-time
persuasion risk
score based on geometric deviation. -
Enhancing Generative Engine Optimization (GEO) Defense: For systems designed for source selection (like generative search engines), implement a defense that specifically targets the pathway identified in the paper: reading persuasive keywords by shallow heads in layers 8–12 to construct the routing feature. This allows the system to filter or down-weight sources whose representations are heavily influenced by these specific, persuasion-driven features.
These improvements would enable AI systems to move from being passively susceptible to context manipulation (like GEO) toward being actively resilient against it by understanding and surgically interrupting the neural pathways responsible for biasing the final output choice.
Abstract
Large language models can answer a factual question correctly, yet switch to an incorrect answer when exposed to persuasive context. Why this happens inside the model remains unclear. In this work, we study this question through causal interventions in controlled multiple-choice tasks and identify that persuasion mainly acts through a small set of mid-layer attention heads, which we call decision heads. These heads do not compute the answer themselves. Instead, they copy the option they attend to, so changing their attention can directly change the model's answer. We find that this attention is controlled by a one-dimensional option-routing feature in the option-token representations. Persuasive text changes this feature, causing the decision heads to attend to the persuasion target instead of the correct option. We further trace the routing feature to shallower attention heads that read persuasive keywords from the input. Together, these results reveal a causal pathway: shallow heads read persuasive keywords and write the routing feature, which redirects the decision heads' attention toward the persuasion target; the decision heads then copy the selected option, leading to the wrong answer. We verify each step through intervention and find the same mechanism across four open-source model families and in a source-selection task derived from Generative Engine Optimization (GEO). Finally, based on this mechanism, we introduce a low-rank update to the decision heads' key projections to reduce their sensitivity to persuasion.
Sources
- How to use and interpret activation patching
- Non-Linear Inference Time Intervention: Improving LLM Truthfulness
- Causal Attribution via Activation Patching
- Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainers
- AtP*: An efficient and scalable method for localizing LLM behaviour to components
- Are You Sure? Challenging LLMs Leads to Performance Drops in The FlipFlop Experiment
- Does Circuit Analysis Interpretability Scale? Evidence from Multiple Choice Capabilities in Chinchilla
- Copy Suppression: Comprehensively Understanding an Attention Head
- The Hydra Effect: Emergent Self-repair in Language Model Computations
- Mass-Editing Memory in a Transformer
- Measuring Opinion Bias and Sycophancy via LLM-based Persuasion
- BrokenMath: A Benchmark for Sycophancy in Theorem Proving with LLMs
- Towards Understanding Sycophancy in Language Models
- Linear Representations of Sentiment in Large Language Models
- Simple synthetic data reduces sycophancy in large language models
- Towards Best Practices of Activation Patching in Language Models: Metrics and Methods
- Flattery in Motion: Benchmarking and Analyzing Sycophancy in Video-LLMs
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection