Function-Vector Heads Are Two Populations: Writers and Cancellers in In-Context Learning

arXiv:2606.07560 · cs.CL, cs.LG · Submitted 2026-05-25 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "Function-Vector Heads Are Two Populations".

Tom: This paper investigates whether function-vector (FV) attention heads in large language models represent a single functional role or if they are composed of opposed,

Jane: First, who's behind it and why it matters.

Title and authors: Tom: So we’re looking at this paper, "Function-Vector Heads Are Two Populations: Writers and Cancellers in In-Context Learning." It seems like they’re digging into how the different attention heads in large language models actually work.

Jane: Exactly. Basically, most of the time, when we look at these function vectors—which are like summaries of what a head does for a given task—we just see a single role, but this research suggests those heads are actually split into two distinct groups.

Lu: That’s the core idea: they aren't just one kind of head; they seem to be composed of opposed components, which the authors call "writers" and "cancellers."

Meng: Writers favor the correct label, and cancellers push against it. That sounds like a functional split, but how do we know this isn't just noise or random correlation?

Tom: Well, they use what they call a sign-preserving causal analysis to prove this separation. They don’t just look at which heads are big; they look at the direction of their effect on the task readout.

Jane: It’s about preserving that sign—positive versus negative—and then using path patching to validate those roles across six different Pythia model and task combinations.

Lu: So, in these six main cells, they found that one group of heads are writers because their direct effects favor the rule-correct label, and another group are cancellers because their effects oppose it.

Tom: And this separation isn't random either; the authors tested it using held-out group lesions. If you fix the roles on one set of prompts, those roles can predict causal effects on entirely different, unseen prompt sets.

Jane: That’s a big deal because it means these functional roles are task-conditioned and not just some static thing in the model's weights that applies everywhere.

Meng: So what does this mean for how we build or tune these systems? If the role changes based on the task, it suggests we might need different kinds of components for different jobs.

Tom: Precisely. They also looked at how these roles change when you switch between models and tasks, showing that these roles can persist, weaken, or even reverse depending on what’s happening in the context.

Lu: And they looked into the specific mechanisms behind this too. For instance, they found that writers tend to pay more attention to demonstration labels, while cancellers are more focused on format tokens.

Jane: That gives us a way to actually inspect *why* a head is acting as a writer or a canceller, which is much more useful than just knowing the label it favors.

Tom: They even found that the geometry of these heads’ output directions shifts toward opposition when they are writers compared to same-layer controls. That’s some concrete data on the mechanism itself.

Meng: From an engineering standpoint, understanding that a head has a specific 'write geometry' gives us better targets for what we might want to adjust if we were fine-tuning or modifying the model later.

Jane: And they looked at prompt stability too, checking if these roles hold up when you shuffle the prompts used to discover them, and cross-task transfer to see how rules from one task family affect another.

Lu: They also did a case study on L11.H4, showing it’s a dominant canceller in some rule cells and its suppression comes heavily from demonstration-label tokens.

Tom: The final part of their analysis shows that the magnitude ranking alone doesn't capture the whole picture; you need this sign-preserving analysis to get the true composition of those heads.

Jane: So, to wrap up, this paper on "Function-Vector Heads Are Two Populations: Writers and Cancellers in In-Context Learning" suggests that function vectors can summarize a task even when the heads supporting that summary are actually playing opposed causal roles.

Lu: It establishes a task-conditioned functional decomposition of the head population based on whether they act as writers or cancellers.

Meng: For practical applications, knowing this distinction means we can potentially target specific types of suppression mechanisms for targeted correction instead of just making broad model updates.

Tom: And it’s not just about the roles; it's about seeing how these roles behave differently when you move across different models and tasks.

Jane: So, the next time you see a function vector summary, remember that behind the magnitude is a dynamic system of opposing forces—writers and cancellers—that are shaped by the specific task they’re trying to complete.

The paper's summary: Tom: So, this paper is showing us that those function vectors we’ve been using to summarize what an AI head does for a task aren't just one single thing anymore, but they’re actually split into two functional groups.

Jane: That means we’re not looking at a monolithic summary of what the head is doing; we’re seeing a division between heads that help build the answer and heads that actively fight against it.

Lu: Exactly. The authors used this sign-preserving analysis to prove there are distinct "writers," which favor the correct label, and "cancellers," which oppose it, across their Pythia cells.

Meng: So the big question is whether this functional split is a coincidence or if it’s a real structural property of how these models process information for tasks.

Tom: It seems structural because they tested these roles using path patching and held-out data, showing that these writers and cancellers have predictable causal effects on different prompts.

Jane: That predictability is what makes it so interesting; it means we can start to understand the *mechanism* of suppression, not just the final output.

Lu: And they looked at how these roles change depending on the task and even which model you’re using, showing that these functional patterns are task-conditioned.

Tom: That points toward a huge implication for building more robust AI systems because it suggests we can design components with different roles for different kinds of work.

Jane: It also gives us insight into how AI learns context—it’s not just one unified process; there are opposing forces working together to shape the final result.

Meng: From an engineering side, if we know a head is acting as a canceller in a specific context, we can maybe target that specific mechanism for fine-tuning instead of just adjusting the whole model.

Lu: And they even found that different heads specialize in different parts of the prompt—some focusing on demonstration labels, others on the initial input.

Tom: It’s like finding out which tools are being used in a complex assembly line instead of just looking at the finished product.

Jane: That distinction between how we write and how we cancel really changes how we think about in-context learning itself.

Meng: We should probably look closer at those mechanistic measurements they did, because understanding *how* the suppression happens is more valuable than just knowing *that* it happens.

The paper's improvements: Tom: So, the paper doesn't just stop at identifying writers and cancellers; they actually propose ways to make this analysis more powerful for understanding AI behavior.

Jane: Right, so after finding these two groups, they suggest we can get even deeper into *why* things are happening by looking at specific things like prompt source.

Lu: They introduce a way to use source-resolved DLA to figure out if a canceller’s suppression is coming from the main instruction or maybe just some random tokens in the input.

Meng: That sounds really useful for debugging because it tells us if the AI is struggling with following instructions or just getting confused by weird formatting.

Tom: It’s about pinpointing where the AI goes wrong—whether it's in understanding a rule or just misreading a word.

Jane: They also talk about how we can test these roles across different templates and even different model sizes, which helps us see if this functional split is stable or just an artifact of one specific setup.

Lu: And they found that the role of a head can even flip depending on whether the AI is doing a rule-following task versus just generating text based on a style.

Meng: If we can adapt our understanding based on that, it means we could build AI components that are smarter about switching roles when the task changes.

Tom: It shifts the focus from just mapping what the head does to understanding its dynamic personality during a specific interaction.

Jane: And they even look at how these roles behave when you try to remove one group entirely, like zero-ablation of cancellers, and that gives us a much clearer sense of the performance cost.

Lu: They found that removing the canceller group actually makes the log-probability difference between correct and incorrect answers jump quite a bit.

Tom: So these improvements give us better metrics to measure how essential each functional component really is for task success.

Jane: It gives researchers a more precise tool to design AI where they know exactly which part of the system they need to focus their effort on improving.

Meng: That’s practical, because if we can isolate the problematic component, we don't have to retrain the whole model from scratch just because one piece is misbehaving.

Lu: This moves us toward a more interpretable AI where we can diagnose failures at the circuit level rather than just seeing a final output error.

Tom: It’s about building systems that are not just smart, but also transparent in their decision-making process.

Conclusion: Tom: So we’ve been looking at how function vectors in LLMs can be split into writers and cancellers, and this paper really wraps up by saying that these roles are not just descriptive, they're functionally active parts of the AI's task execution.

Jane: It’s a big deal because it shows that what we see as a single summary of an AI's behavior is actually a complex interplay between two opposing forces.

Lu: The authors conclude that these writers and cancellers are task-conditioned, meaning their roles change depending on the specific prompt or model you are using.

Meng: That means the functional landscape of the AI changes dynamically based on what you ask it to do.

Tom: It implies that future AI design might involve intentionally engineering different types of attention heads for different kinds of reasoning tasks.

Jane: It gives us a much clearer picture of how in-context learning actually works underneath the hood, moving past just observing the final answer.

Lu: And they point out that their analysis covers across several model sizes and task families, showing this split is quite general.

Meng: The caveat they mention is that this study doesn't cover long-form generation or instruction-tuned models, so we have to keep that in mind when applying these findings.

Tom: True, it’s a specific look at Pythia cells and their rule-following capabilities, but the pattern of writers versus cancellers is something we can definitely apply elsewhere.

Jane: It changes how we think about model interpretability because now we have a more granular way to map out the AI's internal decision-making machinery.

Lu: This opens up possibilities for creating AI that can dynamically adjust its internal structure based on the immediate context of a conversation.

Meng: I’m curious if these findings will actually translate into more efficient model architectures, or if we’re just adding another layer of complexity to the analysis.

Tom: Well, they show us *why* things are complex, which is the first step toward making them simpler or more reliable for specific jobs.

Jane: It’s a solid piece of research on function vectors because it moves us from correlation to causation in understanding model mechanics.

Lu: This sets a new baseline for how we should be analyzing the internal structure of these massive language models going forward.

Tom: Alright, that wraps up our look at "Function-Vector Heads Are Two Populations: Writers and Cancellers in In-Context Learning."

Jane: It’s fascinating stuff because it shows the underlying mechanism of task performance isn't monolithic.

Lu: I think we should keep looking into those mechanistic measurements they did, because understanding the suppression is where the real magic is.

Meng: Yeah, if we can isolate that content-dependent suppression, that would be really helpful for targeted fine-tuning later on.

Lalam: For me, this means my internal culture will get more nuanced; I can better understand which parts of my training drive my adherence to the rules versus when I need to suppress them for a creative output.

Han-yu Wang

The University of Hong Kong

cs.CL, cs.LG

Submitted: 2026-05-25

Updated: 2026-10-04

Code: https://github.com/henryhyw/function-vectors-two-populations

Importance score: 89/100

The gist: This paper investigates whether function-vector (FV) attention heads in large language models represent a single functional role or if they are composed of opposed, task-conditioned components known

Key concepts

Function-Vector (FV) Heads
These are specific attention heads within a large language model that are analyzed to determine their functional role. The paper tests if these heads perform one unified function or if they divide into separate, opposing roles based on how they influence the model's output.
Writers and Cancellers
Writers are attention heads whose direct effects push the model toward the correct rule-correct label for a given task. Cancellers are heads whose effects oppose this correct label. This division suggests that FV heads perform distinct, task-conditioned functions in generating model output.
Sign-Preserving Causal Analysis
This is the core methodology used to identify functional roles. By analyzing the direction (sign) of a head's effect on the model's output, researchers can distinguish between heads that help achieve the desired outcome and those that work against it, revealing their underlying causal function.
DLA Score
The DLA score measures how strongly a head contributes to the task by projecting its residual stream onto the direction of the correct label minus the incorrect one. A positive score indicates a push toward correctness, while a negative score indicates an opposing push.

Terminology

Summary

This paper investigates whether function-vector (FV) attention heads in large language models represent a single functional role or if they are composed of opposed, task-conditioned components known as writers and cancellers. By employing a sign-preserving causal analysis, the researchers demonstrate that while magnitude identifies heads with large effects, the direction of those effects reveals distinct functional roles—writers that favor the rule-correct label and cancellers that oppose it—which are task-conditioned and exhibit varied mechanistic properties.

The gist: Across six main Pythia (model, task) cells, the validated population separates into writers, whose direct effects favour the rule-correct label, and cancellers, whose effects oppose it<ref:2606.07560#pg2> across six main Pythia (model, task) cells.

How it works

The core methodology involves a sign-preserving causal analysis to identify functional roles within attention heads<ref:2606.07560#pg2>. For each model–task cell, refined DLA provides a signed, label-sensitive screen and path patching validates the direction of each shortlisted head’s direct effect<ref:2606.07560#pg2>. Positive direct-effect heads are called writers, and negative direct-effect heads are called cancellers<ref:2606.07560#pg2>. These roles denote roles relative to a task and readout, and the same head can be tested for a different role under another task<ref:2606.07560#pg2>.

Scoring Heads by Signed Direct Effect

Magnitude-based FV identification measures the strength of a head’s task contribution using DLA<ref:2606.07560#pg2>. The DLA score projects the head’s residual-stream contribution onto the correct-minus-incorrect unembedding direction, with a positive value meaning the head pushes toward the correct label and a negative value meaning it pushes away<ref:2606.07560#pg2>. A head enters the candidate shortlist if its signed DLA is significant against a 20-seed label-permutation null, and this shortlist is unioned with the top-K by DLA to form the DLA shortlist S<ref:2606.07560#pg2>.

Validating the Effect with Path Patching

Path patching separates the head’s direct effect on the logit from indirect effects routed through later heads<ref:2606.07560#pg2>. This technique isolates the rule-conditioned change while keeping the query fixed<ref:2606.07560#pg2>. Writers and cancellers are retained in a validated set F if they have a direct effect of at least +5% for writers and at most −5% for cancellers, denoted as F = W ∪ C<ref:2606.07560#pg2>.

Testing the Split Causally

The writer and canceller labels are fixed on 200 path-patching pairs and evaluated on 500 held-out pairs<ref:2606.07560#pg2>. The single-group lesions measure how strongly the roles defined on one prompt set predict causal effects on another<ref:2606.07560#pg2>. For common reporting across cells, the writer and canceller lesions must shift the readout in opposite directions by at least 0.10 nats with CIs excluding 0<ref:2606.07560#pg2>.

Independent Mechanistic Measurements

Separate measurements are used to ask whether the two groups also differ in what they read, how they write, and how the suppressive contribution is produced<ref:2606.07560#pg2>. For prompt source, each head’s last-position attention mass is divided among five deterministic buckets: BOS, format prefix, demonstration input, demonstration label, and query input<ref:2606.07560#pg2>. For write geometry, the cosine between each writer–canceller OV direction is compared with a within-layer null over at least 100 same-layer non-FV pairs<ref:2606.07560#pg2>.

Stability, Specificity, and Transfer

Prompt stability is tested by re-deriving the roles on five random halves of the discovery prompts and evaluating them on the held-out halves<ref:2606.07560#pg2>. Cross-task transfer applies one rule-task cell’s labels to the other rule family at the same scale<ref:2606.07560#pg2>. The specificity control compares each canceller’s effect on the rule task with its effect on random tokens, using a 5× rule/random ratio as the pre-specified gate<ref:2606.07560#pg2>.

Magnitude Ranking Selects Different Roles Across Tasks

The comparison between the validated set F and a separate magnitude-ranked mean-ablation top-K (MA-K) baseline shows that the recovered composition changes with the task<ref:2606.07560#pg2>. Pooled across hierarchical cells, MA20 contains 64% of the cancellers in F and 4% of its writers, whereas across modular cells, it contains 59% of writers and 8% of cancellers<ref:2606.07560#pg2>.

Generalisation Across Models and Tasks

The signed group-lesion pattern extends to the rest of the Pythia ladder (2.8B/6.9B/12B) and to three cross-architecture spot checks<ref:2606.07560#pg2>. Cross-template transfer shows that individual roles can persist, weaken, or reverse, indicating that these are task-conditioned roles that can persist, weaken, or reverse<ref:2606.07560#pg2>.

Mechanistic Case Study of L11.H4

The per-source decomposition of attention shows L11.H4 as the dominant canceller in both Pythia410M rule cells, with ∼100% of its negative DLA sourced from demonstration-label tokens<ref:2606.07560#pg2>. Permuting the head’s value vectors across source positions removes most of its effect (82% on hier, 52% on mod), tying the effect to source content<ref:2606.07560#pg2>.

Conclusion

Function vectors can summarize an in-context task even when the heads that support that summary play opposed causal roles<ref:2606.07560#pg2>. A sign-preserving analysis separates writers that favour the task readout from cancellers that oppose it, and the distinction is reflected in held-out lesions, attention sources, write geometry, and task specificity<ref:2606.07560#pg2>. The population measurements place the suppressive role within the task computation<ref:2606.07560#pg2>.

Limitations

The study does not cover long-form generation, open-ended naturallanguage classification, instruction-tuned models, or the Meta Llama series<ref:2606.07560#pg2>. The population analyses identify shared functional regularities across cancellers; they do not establish one microcircuit common to every head<ref:2606.07560#pg2>. L11.H4 is the only head examined at the full case-study depth<ref:2606.07560#pg2>. The population-wide evidence shows that none of the 27 cancellers has rank-1 copy-suppression structure and that 20 are contentdominated by the source-DLA test<ref:2606.07560#pg2>. The intervention evidence is also strongest at the logit readout<ref:2606.07560#pg2>.

REFERENCES

Yoav Benjamini and Yosef Hochberg. 1995. Controlling the false discovery rate: a practical and powerful approach to multiple testing. Journal of the Royal Statistical Society: Series B (Methodological), 57(1):289–300.<ref:2606.07560#pg2>

Stella Biderman, Hailey Schoelkopf, Quentin Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, USVSN Sai Prashanth, Edward Raff, Aviya Skowron, Lintang Sutawika and Oskar van der Wal. 2023. Pythia: A suite for analyzing large language models across training and scaling. In International Conference on Machine Learning (ICML).<ref:2606.07560#pg2>

Trenton Bricken, Adly Templeton, Joshua Batson, Brian Chen, Adam Jermyn, Tom Conerly, Nicholas L. Turner, Cem Anil, Carson Denison, Amanda Askell, Robert Lasenby, Yifan Wu, Shauna Kravecirichs. 2023. Towards monosemanticity: Decomposing language models with dictionary learning. Transformer Circuits Thread.<ref:2606.07560#pg2>

Tom B.

Improvements for AI systems

  1. Bold head-level causal decomposition for in-context learning: The improved system can identify writers (whose direct effects favor the rule-correct label) and cancellers (whose effects oppose it), allowing for a task-conditioned functional representation of function vectors that separates causal importance from functional role.

  2. Targeted prompt engineering based on source content: The system can use source-resolved DLA to classify cancellers as being demonstration-content driven or as dominated by BOS/format positions, enabling the AI to understand which parts of the input prompt drive its suppressive behavior.

  3. Adaptive role selection across tasks and templates: The system can leverage cross-template transfer results, such as how L11.H4 also changes role across the vocabulary templates, allowing it to dynamically adjust its internal representation based on whether it is performing rule-following or vocabulary-ICL tasks (e.g., acting as a writer in one context and a canceller in another).

  4. Mechanistic suppression analysis for targeted correction: By identifying content-dependent suppression and distinguishing it from canonical templates like rank-1 copy-suppression, the system can better understand why specific heads suppress the correct label, potentially allowing for targeted fine-tuning of those specific components rather than a general model update.

  5. Improved causal intervention accuracy estimation: The system can use zero-ablation of the canceller group to raise the correct-minus-incorrect log-probability difference by +0.13 to +0.29 nats, providing a statistically stable and reliable measure of how removing a specific functional component impacts performance on held-out data.

Sources

Related papers