MoRFI: Monotonic Sparse Autoencoder Feature Identification

summary

Video file (mp4)

The gist

MoRFI: Monotonic Sparse Autoencoder Feature Identification Abstract Large language models (LLMs) acquire most of their factual knowledge during the pre-training stage, through next token prediction.

In short

The discussion of 'MoRFI: Monotonic Sparse Autoencoder Feature Identification' explores a method for identifying internal model features responsible for hallucinations after fine-tuning. The hosts conclude that these features act as gatekeepers to existing knowledge, and restoring accuracy by steering a single feature offers a path toward more reliable AI systems.

Key concepts

Monotonic Sparse Autoencoder
This is the core methodology used to break down high-dimensional model states into thousands of cleaner, individual features. MoRFI then tracks how these specific features change across a spectrum of fine-tuning conditions to find a consistent, causal signal.
Hallucination Mechanism
The paper proposes that when large language models are fine-tuned on new facts, they don't forget what they knew; instead, their internal wiring gets disrupted. MoRFI aims to find the specific neurons or features responsible for this breakdown.
Feature Identification/Steering
This is the process of locating specific internal features and then intervening on them. By manipulating a single identified feature, researchers can restore a significant chunk of lost accuracy, suggesting that knowledge is being blocked rather than erased.

Terminology used across episodes

This episode discusses

The paper

MoRFI: Monotonic Sparse Autoencoder Feature Identification · Read on arXiv

Dimitris Dimakopoulos, Shay B. Cohen, Ioannis Konstas

University of Edinburgh · Heriot-Watt University

Large language models (LLMs) acquire most of their factual knowledge during the pre-training stage, through next token prediction. Subsequent stages of post-training often introduce new facts outwith the parametric knowledge, giving rise to hallucinations. While it has been demonstrated that supervised fine-tuning (SFT) on new knowledge may exacerbate the problem, the underlying mechanisms are still poorly understood. We conduct a controlled fine-tuning experiment, focusing on closed-book QA, and find latent directions that causally contribute to hallucinations. Specifically, we fine-tune Llama 3.1 8B, Gemma 2 9B and Mistral 7B v03 on seven distinct single QA datasets, controlling for the percentage of new knowledge and number of training epochs. By measuring performance on the test set, we validate that incrementally introducing new knowledge increases hallucinations, with the effect being more pronounced with prolonged training. We leverage pre-trained sparse autoencoders (SAEs) to analyze residual stream activations across various checkpoints for each model and propose Monotonic Relationship Feature Identification (MoRFI) for capturing causally relevant latents. MoRFI filters SAE features that respond monotonically to controlled fine-tuning data mixtures of a target property. Our findings show that exposure to unknown facts disrupts the model's ability to retrieve stored knowledge along a set of directions in the residual stream. Our pipeline reliably discovers them across distinct models, recovering knowledge through single-latent interventions.

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "MoRFI: Monotonic Sparse Autoencoder Feature Identification".

Jane: The paper was written by Dimitris Dimakopoulos, Shay B. Cohen and Ioannis Konstas from University of Edinburgh and Heriot-Watt University.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title: Tom: Welcome back to the show, everyone. Today we're digging into a fresh preprint that just hit arXiv, and the title alone tells you it's going to be a dense one: "MoRFI: Monotonic Sparse Autoencoder Feature Identification." Jane, I'll be honest, my first reaction was, that's a mouthful. What's actually going on here?

Jane: It is a mouthful, Tom, but the idea underneath is really elegant. The paper is about why large language models start hallucinating when you fine-tune them on new facts. The authors have a hunch that it's not that the model forgets what it knew, but that something in its internal wiring gets disrupted. MoRFI is their method for finding those specific wires.

Tom: So we're talking about looking inside the black box, finding the exact neurons or features that are responsible for this breakdown?

Jane: Exactly. They use something called sparse autoencoders, which are like a magnifying glass for the model's internal activations. It breaks down the messy, high-dimensional state into thousands of cleaner, individual features. MoRFI then looks at how those features change as they fine-tune the model on data that is increasingly full of facts the model doesn't know.

Lu: And that's the clever part, Jane. They're not just looking at one model state versus another. They're creating a whole spectrum of models, from one fine-tuned on zero unknown facts to one fine-tuned on one hundred percent unknown facts. Then they track which features respond in a steady, monotonic way across that whole spectrum. It’s a much more rigorous way to find a causal signal than just comparing two snapshots.

Meng: So it's a feature selection algorithm. But from an engineering standpoint, what does "monotonic" buy us? Why not just look at the features that change the most between the best and worst performing models?

Jane: Great question, Meng. The problem with just comparing two points is that you might catch a feature that spiked for some random reason, like overfitting to a specific prompt. By requiring a consistent, monotonic trend across seven different data mixtures, MoRFI filters out those transient blips. It's looking for a feature that is steadily increasing or decreasing as the problem gets worse, which is a much stronger signal that it's actually tied to the hallucination mechanism.

Tom: And the payoff here is huge, right? They're not just identifying these features for fun. They're actually steering the model with them, and they're recovering a significant chunk of the lost accuracy. We're talking about a potential on-the-fly fix for hallucinations.

Lu: Precisely, Tom. The fact that they can intervene on a single feature and restore performance suggests that the knowledge isn't erased; it's just being blocked. That's a profound finding for how we think about model editing and safety.

Tom: So we've got a method, we've got a potential fix. But I'm already wondering, how do they know these specific features are the right ones and not just some arbitrary artifact? That's the question I want to dig into next.

Summary: Tom: So we've established that "MoRFI: Monotonic Sparse Autoencoder Feature Identification" is about finding the internal culprits for hallucinations after fine-tuning. But Jane, I'm still stuck on the validation. How do they prove these features are actually the ones that matter?

Jane: That's the most important part, Tom. They run a control experiment. They pick a separate group of features that show no trend at all across the fine-tuning spectrum. Then they try steering the model with those control features, and guess what happens?

Tom: Let me guess, nothing good?

Jane: Exactly. Steering with the control features either does nothing or actively hurts performance. But when they steer with the top features identified by MoRFI, they see substantial accuracy gains. That contrast is what proves the signal they found is real and not just random noise in the residual stream.

Meng: And the numbers back that up. On Llama three point one 8B, they got a relative accuracy gain of over thirty-three percent on the dev set by steering a single, negatively-trending feature. That's a massive jump from a baseline of seventeen point eight percent. The control group, meanwhile, gave them basically nothing.

Lu: What's even more interesting to me is the asymmetry they found. Suppressing features that increased during fine-tuning was much more effective than amplifying features that decreased. It’s like the model is shouting the wrong thing, and you just need to turn the volume down, rather than trying to whisper the right thing louder.

Jane: Right, and that leads to their knowledge recovery metric. They showed that sixty-nine to eighty-five percent of the facts they recovered were facts the original, healthy model knew. That's a really strong piece of evidence that these features are gatekeepers to existing knowledge, not a path to learning new facts.

Tom: So the model hasn't forgotten, it's just being blocked from accessing what it knows. That's a wild thought. It's like the fine-tuning process puts a cork in the bottle, and MoRFI finds the exact cork to pull out.

Meng: But I have to ask, how practical is this? We're talking about steering a model at inference time. Is this something we could do in production, or is it just a research lab trick?

Jane: That's a fair question. The paper shows it works across three different model families, which is a good sign for generality. But the process of finding these features requires training multiple checkpoints and running a bootstrap analysis, which is computationally heavy. It's not a real-time fix yet, but it's a powerful diagnostic tool and a proof of concept for a whole new class of interventions.

Tom: So it's a roadmap for a future fix, not the fix itself. I can get behind that. But what I really want to know is, what does this mean for the bigger picture? We're talking about making models more reliable, but what are the broader implications? Let's take a quick break and we'll get into that.

Improvements: Tom: Welcome back. We're deep in the weeds of "MoRFI: Monotonic Sparse Autoencoder Feature Identification," and we've seen how it finds and validates these knowledge-blocking features. Jane, I want to zoom out now. What's the actual improvement this paper offers over what was already out there?

Jane: The biggest improvement is the methodology itself. Previous work often compared a model before and after fine-tuning, which can pick up on all sorts of unrelated changes. MoRFI instead looks at a controlled gradient of change. By varying the percentage of unknown facts from zero to one hundred percent, they're isolating the effect of new knowledge specifically, rather than just the effect of training in general.

Lu: And that's a crucial distinction, Jane. It's the difference between finding a correlation and finding a cause. The monotonic trend requirement is what gives them the confidence to say these features are causally linked to the hallucination behavior, not just correlated with it. It's a much more rigorous standard for mechanistic interpretability.

Meng: From a practical standpoint, the improvement is in the efficiency of the intervention. They found that steering with a single latent can outperform steering with a composite direction that represents the entire shift in activation space. That's a huge deal. It means the signal is incredibly sparse, and you don't need to make broad, blunt changes to the model to fix a specific problem. You can make a surgical strike.

Tom: Surgical strike, I like that. So instead of trying to rewind the whole model to its pre-fine-tuned state, you just nudge one little lever and the knowledge comes flooding back. That's incredibly efficient.

Jane: Exactly. And it also gives us a new tool for understanding model behavior. These features aren't just for fixing hallucinations. The authors note that the features they found are interpretable, like one that activates on geographic locations. This means we can start to map out where different types of knowledge live in the model and how they're connected.

Lu: The implications for safety are huge, too. If we can identify features that control access to knowledge, we might be able to identify features that control other behaviors, like following harmful instructions or being deceptive. This gives us a much more precise toolkit for auditing and controlling model behavior.

Meng: So it's not just a fix, it's a diagnostic tool and a map. That's a pretty powerful combination. But I'm still curious about the limitations. The paper uses closed-book QA, which is a pretty narrow task. How far does this generalize?

Jane: That's the open question, Meng. The paper is a proof of concept on a specific task, but the MoRFI algorithm itself is task-agnostic. It's designed to find features that respond monotonically to any controlled variable. So the potential is there to apply it to other behaviors, but that's future work.

Tom: So we've got a powerful new method, a proof of concept, and a whole lot of potential. I think it's time we started wrapping our heads around what this all means for the future. Let's head to the conclusion.

Conclusion: Tom: Well, we've spent a good chunk of the show unpacking "MoRFI: Monotonic Sparse Autoencoder Feature Identification," and I think it's safe to say this is one of those papers that makes you see the field a little differently. Jane, can you give us the one-minute version for anyone just tuning in?

Jane: Sure, Tom. The paper tackles the mystery of why fine-tuning on new facts causes hallucinations. Their answer is that it's not about forgetting, it's about blocking access. They developed a method called MoRFI to find the specific internal features that are responsible for this blockage, and they showed that by carefully steering just one of those features, you can restore a large portion of the model's lost knowledge.

Lu: And the key to their success was the rigorous methodology. By looking for features with a monotonic trend across a spectrum of fine-tuning conditions, they were able to separate causal signals from noise. This sets a new standard for how we do feature identification in mechanistic interpretability.

Meng: I'm still impressed by the practical angle. The fact that a single-latent intervention can outperform a broad, composite one suggests that future model editing could be far more surgical and efficient than we thought possible. It’s a promising direction for building more reliable and controllable systems.

Tom: And on a broader level, this paper gives us a new way to think about knowledge in these models. It's not just stored in a big blob; it's gated and managed by specific circuits. Understanding those circuits is going to be key to making eye we can truly trust.

Lalam: It also opens a cultural door, Tom. If we can reliably prevent hallucinations, we can start using these models in domains where accuracy is paramount, like education, historical preservation, and scientific communication. It moves eye from a creative toy to a dependable collaborator, enriching how we share and verify knowledge as a society.

Jane: That's a beautiful way to put it, Lalam. So, we've said our piece on MoRFI. It's a dense paper with a powerful idea, and we'll be keeping an eye on where this line of research goes. Thanks for joining us, everyone. We'll see you on the next one.

More episodes

← Home