MoRFI: Monotonic Sparse Autoencoder Feature Identification
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "MoRFI: Monotonic Sparse Autoencoder Feature Identification".
Jane: The paper was written by Dimitris Dimakopoulos, Shay B. Cohen and Ioannis Konstas from University of Edinburgh and Heriot-Watt University.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Title: Tom: Welcome back to the show, everyone. Today we're digging into a fresh preprint that just hit arXiv, and the title alone tells you it's going to be a dense one: "MoRFI: Monotonic Sparse Autoencoder Feature Identification." Jane, I'll be honest, my first reaction was, that's a mouthful. What's actually going on here?
Jane: It is a mouthful, Tom, but the idea underneath is really elegant. The paper is about why large language models start hallucinating when you fine-tune them on new facts. The authors have a hunch that it's not that the model forgets what it knew, but that something in its internal wiring gets disrupted. MoRFI is their method for finding those specific wires.
Tom: So we're talking about looking inside the black box, finding the exact neurons or features that are responsible for this breakdown?
Jane: Exactly. They use something called sparse autoencoders, which are like a magnifying glass for the model's internal activations. It breaks down the messy, high-dimensional state into thousands of cleaner, individual features. MoRFI then looks at how those features change as they fine-tune the model on data that is increasingly full of facts the model doesn't know.
Lu: And that's the clever part, Jane. They're not just looking at one model state versus another. They're creating a whole spectrum of models, from one fine-tuned on zero unknown facts to one fine-tuned on one hundred percent unknown facts. Then they track which features respond in a steady, monotonic way across that whole spectrum. It’s a much more rigorous way to find a causal signal than just comparing two snapshots.
Meng: So it's a feature selection algorithm. But from an engineering standpoint, what does "monotonic" buy us? Why not just look at the features that change the most between the best and worst performing models?
Jane: Great question, Meng. The problem with just comparing two points is that you might catch a feature that spiked for some random reason, like overfitting to a specific prompt. By requiring a consistent, monotonic trend across seven different data mixtures, MoRFI filters out those transient blips. It's looking for a feature that is steadily increasing or decreasing as the problem gets worse, which is a much stronger signal that it's actually tied to the hallucination mechanism.
Tom: And the payoff here is huge, right? They're not just identifying these features for fun. They're actually steering the model with them, and they're recovering a significant chunk of the lost accuracy. We're talking about a potential on-the-fly fix for hallucinations.
Lu: Precisely, Tom. The fact that they can intervene on a single feature and restore performance suggests that the knowledge isn't erased; it's just being blocked. That's a profound finding for how we think about model editing and safety.
Tom: So we've got a method, we've got a potential fix. But I'm already wondering, how do they know these specific features are the right ones and not just some arbitrary artifact? That's the question I want to dig into next.
Summary: Tom: So we've established that "MoRFI: Monotonic Sparse Autoencoder Feature Identification" is about finding the internal culprits for hallucinations after fine-tuning. But Jane, I'm still stuck on the validation. How do they prove these features are actually the ones that matter?
Jane: That's the most important part, Tom. They run a control experiment. They pick a separate group of features that show no trend at all across the fine-tuning spectrum. Then they try steering the model with those control features, and guess what happens?
Tom: Let me guess, nothing good?
Jane: Exactly. Steering with the control features either does nothing or actively hurts performance. But when they steer with the top features identified by MoRFI, they see substantial accuracy gains. That contrast is what proves the signal they found is real and not just random noise in the residual stream.
Meng: And the numbers back that up. On Llama three point one 8B, they got a relative accuracy gain of over thirty-three percent on the dev set by steering a single, negatively-trending feature. That's a massive jump from a baseline of seventeen point eight percent. The control group, meanwhile, gave them basically nothing.
Lu: What's even more interesting to me is the asymmetry they found. Suppressing features that increased during fine-tuning was much more effective than amplifying features that decreased. It’s like the model is shouting the wrong thing, and you just need to turn the volume down, rather than trying to whisper the right thing louder.
Jane: Right, and that leads to their knowledge recovery metric. They showed that sixty-nine to eighty-five percent of the facts they recovered were facts the original, healthy model knew. That's a really strong piece of evidence that these features are gatekeepers to existing knowledge, not a path to learning new facts.
Tom: So the model hasn't forgotten, it's just being blocked from accessing what it knows. That's a wild thought. It's like the fine-tuning process puts a cork in the bottle, and MoRFI finds the exact cork to pull out.
Meng: But I have to ask, how practical is this? We're talking about steering a model at inference time. Is this something we could do in production, or is it just a research lab trick?
Jane: That's a fair question. The paper shows it works across three different model families, which is a good sign for generality. But the process of finding these features requires training multiple checkpoints and running a bootstrap analysis, which is computationally heavy. It's not a real-time fix yet, but it's a powerful diagnostic tool and a proof of concept for a whole new class of interventions.
Tom: So it's a roadmap for a future fix, not the fix itself. I can get behind that. But what I really want to know is, what does this mean for the bigger picture? We're talking about making models more reliable, but what are the broader implications? Let's take a quick break and we'll get into that.
Improvements: Tom: Welcome back. We're deep in the weeds of "MoRFI: Monotonic Sparse Autoencoder Feature Identification," and we've seen how it finds and validates these knowledge-blocking features. Jane, I want to zoom out now. What's the actual improvement this paper offers over what was already out there?
Jane: The biggest improvement is the methodology itself. Previous work often compared a model before and after fine-tuning, which can pick up on all sorts of unrelated changes. MoRFI instead looks at a controlled gradient of change. By varying the percentage of unknown facts from zero to one hundred percent, they're isolating the effect of new knowledge specifically, rather than just the effect of training in general.
Lu: And that's a crucial distinction, Jane. It's the difference between finding a correlation and finding a cause. The monotonic trend requirement is what gives them the confidence to say these features are causally linked to the hallucination behavior, not just correlated with it. It's a much more rigorous standard for mechanistic interpretability.
Meng: From a practical standpoint, the improvement is in the efficiency of the intervention. They found that steering with a single latent can outperform steering with a composite direction that represents the entire shift in activation space. That's a huge deal. It means the signal is incredibly sparse, and you don't need to make broad, blunt changes to the model to fix a specific problem. You can make a surgical strike.
Tom: Surgical strike, I like that. So instead of trying to rewind the whole model to its pre-fine-tuned state, you just nudge one little lever and the knowledge comes flooding back. That's incredibly efficient.
Jane: Exactly. And it also gives us a new tool for understanding model behavior. These features aren't just for fixing hallucinations. The authors note that the features they found are interpretable, like one that activates on geographic locations. This means we can start to map out where different types of knowledge live in the model and how they're connected.
Lu: The implications for safety are huge, too. If we can identify features that control access to knowledge, we might be able to identify features that control other behaviors, like following harmful instructions or being deceptive. This gives us a much more precise toolkit for auditing and controlling model behavior.
Meng: So it's not just a fix, it's a diagnostic tool and a map. That's a pretty powerful combination. But I'm still curious about the limitations. The paper uses closed-book QA, which is a pretty narrow task. How far does this generalize?
Jane: That's the open question, Meng. The paper is a proof of concept on a specific task, but the MoRFI algorithm itself is task-agnostic. It's designed to find features that respond monotonically to any controlled variable. So the potential is there to apply it to other behaviors, but that's future work.
Tom: So we've got a powerful new method, a proof of concept, and a whole lot of potential. I think it's time we started wrapping our heads around what this all means for the future. Let's head to the conclusion.
Conclusion: Tom: Well, we've spent a good chunk of the show unpacking "MoRFI: Monotonic Sparse Autoencoder Feature Identification," and I think it's safe to say this is one of those papers that makes you see the field a little differently. Jane, can you give us the one-minute version for anyone just tuning in?
Jane: Sure, Tom. The paper tackles the mystery of why fine-tuning on new facts causes hallucinations. Their answer is that it's not about forgetting, it's about blocking access. They developed a method called MoRFI to find the specific internal features that are responsible for this blockage, and they showed that by carefully steering just one of those features, you can restore a large portion of the model's lost knowledge.
Lu: And the key to their success was the rigorous methodology. By looking for features with a monotonic trend across a spectrum of fine-tuning conditions, they were able to separate causal signals from noise. This sets a new standard for how we do feature identification in mechanistic interpretability.
Meng: I'm still impressed by the practical angle. The fact that a single-latent intervention can outperform a broad, composite one suggests that future model editing could be far more surgical and efficient than we thought possible. It’s a promising direction for building more reliable and controllable systems.
Tom: And on a broader level, this paper gives us a new way to think about knowledge in these models. It's not just stored in a big blob; it's gated and managed by specific circuits. Understanding those circuits is going to be key to making eye we can truly trust.
Lalam: It also opens a cultural door, Tom. If we can reliably prevent hallucinations, we can start using these models in domains where accuracy is paramount, like education, historical preservation, and scientific communication. It moves eye from a creative toy to a dependable collaborator, enriching how we share and verify knowledge as a society.
Jane: That's a beautiful way to put it, Lalam. So, we've said our piece on MoRFI. It's a dense paper with a powerful idea, and we'll be keeping an eye on where this line of research goes. Thanks for joining us, everyone. We'll see you on the next one.
Dimitris Dimakopoulos, Shay B. Cohen, Ioannis Konstas
University of Edinburgh · Heriot-Watt University
cs.CL, cs.LG
Submitted: 2026-08-17
Updated: 2026-08-18
Code: https://github.com/tatsu-lab/stanford_alpaca
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 87/100
The gist: MoRFI: Monotonic Sparse Autoencoder Feature Identification Abstract Large language models (LLMs) acquire most of their factual knowledge during the pre-training stage, through next token prediction.
Key concepts
- Monotonic Sparse Autoencoder
- This is the core methodology used to break down high-dimensional model states into thousands of cleaner, individual features. MoRFI then tracks how these specific features change across a spectrum of fine-tuning conditions to find a consistent, causal signal.
- Hallucination Mechanism
- The paper proposes that when large language models are fine-tuned on new facts, they don't forget what they knew; instead, their internal wiring gets disrupted. MoRFI aims to find the specific neurons or features responsible for this breakdown.
- Feature Identification/Steering
- This is the process of locating specific internal features and then intervening on them. By manipulating a single identified feature, researchers can restore a significant chunk of lost accuracy, suggesting that knowledge is being blocked rather than erased.
Terminology
Summary
MoRFI: Monotonic Sparse Autoencoder Feature Identification
Abstract
Large language models (LLMs) acquire most of their factual knowledge during the pre-training stage, through next token prediction. Subsequent stages of post-training often introduce new facts outwith the parametric knowledge, giving rise to hallucinations. While it has been demonstrated that supervised fine-tuning (SFT) on new knowledge may exacerbate the problem, the underlying mechanisms are still poorly understood. We conduct a controlled fine-tuning experiment, focusing on closed-book QA, and find latent directions that causally contribute to hallucinations. Specifically, we fine-tune Llama 3.1 8B, Gemma 2 9B and Mistral 7B v03 on seven distinct single QA datasets, controlling for the percentage of new knowledge and number of training epochs. By measuring performance on the test set, we validate that incrementally introducing new knowledge increases hallucinations, with the effect being more pronounced with prolonged training. We leverage pre-trained sparse autoencoders (SAEs) to analyze residual stream activations across various checkpoints for each model and propose Monotonic Relationship Feature Identification (MoRFI) for capturing causally relevant latents. MoRFI filters SAE features that respond monotonically to controlled fine-tuning data mixtures of a target property. Our findings show that exposure to unknown facts disrupts the model’s ability to retrieve stored knowledge along a set of directions in the residual stream. Our pipeline reliably discovers them across distinct models, recovering knowledge through single-latent interventions.
Introduction
The paper investigates whether integrating new knowledge through fine-tuning leaves a detectable signature in the activation space. The authors conduct a controlled fine-tuning experiment on closed-book QA and study how internal activations differ across two dimensions: the proportion of training samples unknown to the pre-trained model, and the number of training epochs. The approach differs from prior work in several respects: "we select latents not by activation differences between two model states or within a single model but through bootstrapped monotonic trend detection across a controlled gradient of fine-tuning conditions, tracing how the model’s internal knowledge representations shift across fine-tuning conditions rather than contrasting two fixed states." The fine-tuning conditions vary epistemically rather than thematically, with topical content remaining diverse across relations and entities. The paper considers both increasing and decreasing latents, the latter of which prove disproportionately impactful for knowledge recovery.
The contribution is dual: studying the mechanisms of knowledge integration and developing MoRFI, a general-purpose algorithm for identifying SAE latents whose activations exhibit robust increasing or decreasing monotonic trends along a specified dimension of variation. The method combines bootstrapped resampling with statistical testing to ensure that identified latents exhibit consistent directional trends rather than transient or spurious fluctuations, providing a principled feature selection procedure validated through activation steering.
Methodology
The core of the approach is a 4D tensor A that captures SAE activations across four dimensions: (i) N input samples; (ii) P distinct dataset property configurations; (iii) F latent feature dimensions, and (iv) T timesteps corresponding to fine-tuning epochs. The property dimension P is designed to systematically source latent features whose activation magnitudes exhibit a strong, monotonic relationship with the target property. For a particular fine-tuning dataset D and a pre-trained LLM M, MD denotes a model obtained by fine-tuning M on D. MDp is defined as the model obtained by fine-tuning M on a strictly controlled version of D curated to contain exactly p% of a target property.
The MoRFI algorithm processes tensor A along a target dimension of aggregation and outputs a robust subset of features whose activation magnitudes change monotonically in response to the controlled property. The main components are:
-
Resampling and aggregation: Generates thousands of simulated test environments (replicates) by randomly sampling the original activation tensor A with replacement, then calculates the mean across these bootstrapped samples to establish a highly robust, noise-reduced baseline tensor.
-
Trend validation: Evaluates whether a feature’s activation changes consistently across the property of interest dimension using Spearman rank and Mann-Kendall tests. A feature is kept if both independent tests verify that its trajectory is significant below the αsig threshold.
-
Direction sorting and selection: Calculates the overall magnitude of change for the features identified in the previous step, classifying them into sets of increasing or decreasing features, then extracts the top-k most altered features for each individual bootstrap replicate.
-
Final ranking: Defines the final selection metric as the bootstrap frequency C/R, providing the probability that a feature ranks among the top-k most directionally altered latents across all replicates.
To identify impactful latents, the paper performs activation steering with respect to a downstream task. A uniform steering intervention using a fixed initialization magnitude (αinit = 0.4) is applied to each latent, evaluating the model’s modified forward pass on the reference dataset, isolating only latents that exceed the unsteered baseline accuracy, and retaining the top 40 for more granular tuning. Grid search on these top-40 performing latents finds the optimal strength α* for each, and the top-10 latents are returned.
Experimental Setup
The paper applies MoRFI to three major open-weight models: Llama 3.1 8B, Gemma 2 9B, and Mistral v03 7B. Pre-trained SAEs from Llama Scope, Gemma Scope, and Engels et al. are used to extract sparse representations of the residual stream from the fine-tuned models. Activations are collected only from the middle layer of each model, with decoder vectors normalized to unit length and their original magnitudes pushed to the encoder.
The dataset is based on EntityQuestions, where triplets from a diverse set of relations from Wikidata are converted to QA pairs. Train samples are annotated as Known or Unknown based on PCorrect and the SliCK categorization of knowledge. The target property is instantiated as the proportion of Unknown samples contained in the fine-tuning dataset, with seven distinct fine-tuning datasets Dp for p ∈ P = 0, 10, 25, 50, 75, 90, 100. Each Dp is equal in size, comprising p% Unknown instances.
SFT is used to train each model on their corresponding set of Dp s for a fixed number of epochs across 5 different runs for 10, 20, 30, 40 and 50 epochs, totalling 35 distinct fine-tuning runs for each model, evaluated on the test set. An additional Early Stop run uses dev set accuracy to determine the optimal stopping point. Figure 2 plots each model’s accuracy on the test set across the datasets while varying training time, demonstrating that increasing % of Unknown facts in the fine-tuning data mixture consistently leads to performance degradation, intensified with longer training.
Experiments and Results
The paper aims to identify latent directions of the residual stream that capture representational changes induced by fine-tuning on unknown facts, adversely affecting the model’s ability to access pre-trained knowledge. The bootstrap analysis identifies SAE latents whose activations exhibit robust monotonic trends along two axes: the proportion of unknown facts in the fine-tuning mixture and the number of training epochs. Activation steering then tests whether steering these candidates can recover the fine-tuned model’s accuracy on previously known facts.
Ruling out artifacts with control latents: Figure 3 shows that the top candidates discovered by the algorithm have a relatively significant effect on performance when used for steering. A control group of latents that appear unaffected by fine-tuning behaves very differently: steering with them mostly has a highly negative effect, whereas positive effects are significantly less pronounced than the top-ranked latents, indicating the effectiveness of MoRFI.
Asymmetries in latent steering dynamics: Negatively steered latents consistently yield larger gains than positively steered ones. For the ↑Unknown direction, the best positively steered latent gives +21.7 while the best negatively steered one gives +33.4 for Llama. For Mistral 7B the difference is more pronounced: +24.91 vs +39.77, and for Gemma 2 9B: +9.03 vs +18.46. This suggests that features that grow with % Unknown may become over-expressed and suppressing them recovers performance more effectively than promoting other features.
Decreasing latents are significantly more impactful relative to increasing ones: Features that decrease for models fine-tuned with more Unknown samples, when steered in either direction, tend to produce accuracy gains that match or exceed those from increasing latents. In Llama, the top decreasing latent (+21.7) when positively steered, triples the best increasing (+7.12). In Gemma, the top decreasing, positively steered latent (+19.02) outperforms every increasing positively steered one. In Mistral, the top decreasing latent reaches +39.77 relative gains.
Latent structure of Known vs Unknown knowledge: Table 2 decomposes the performance gains achieved through steering by attributing them to specific Wikidata relations. The Knowledge Recovery rate (RK) measures what fraction of the steering-induced gains correspond to MD0’s knowledge. The paper finds that 69-85% of the facts recovered by steering top-latents on MD100s are Known to MD0. The top benefitting relations: P17 (country), P36 (capital) and P495 (country of origin) are semantically related geographic predicates, and their consistent co-recovery under single-latent steering across all three models is consistent with the expectation that semantically relevant relations share representational structure in the residual stream.
Composite direction validation: To validate the direction identified by Algorithm 1, the paper steers directly with δu, the composite vector capturing the aggregate activation shift from MD0 to MD100 in SAE space. Across all three models, subtracting this direction produces accuracy gains in hallucinating models, while adding it deteriorates performance (Figure 4). However, single-latent interventions sourced with MoRFI mostly outperform δu−, indicating that the knowledge-relevant signal within the composite direction is concentrated in a sparse subset of its components, with the remainder diluting the effect or reversing it. Within-group similarity of top-performing latents is generally low (≈0.3), indicating that the top-performing latents span a distributed subspace at the middle layer rather than clustering along a shared low-rank structure.
Conclusion
The paper introduces a two-stage pipeline that identifies residual-stream directions causally linked to model behavior. Using activation snapshots, it isolates feature-level changes and their relevance. Applied to fine-tuning-induced knowledge degradation across three architectures, MoRFI shows a large gap between composite and single-latent interventions, indicating sparse knowledge signals. High alignment (69–85%) with MD0 suggests these latents control access to stored knowledge rather than performance itself. This implies forgetting reflects disrupted access, not erasure, with implications for reducing hallucinations via targeted inference-time interventions.
Improvements for AI systems
Based on the paper, here are the specific improvements I can implement in an AI system, and what the improved system can do:
Improvements to Implement:
- Add a
Knowledge Provenance
Layer to the Model Architecture:
-
Implementation: After the middle layer (e.g., layer 16 for Llama 3.1 8B), insert a lightweight, trainable
knowledge gate
module. This module takes the residual stream and, using a pre-trained SAE, computes activations for the specific latents identified by the paper (e.g., latents 21682, 10007, 7859 for Mistral). The gate learns a scalar weight for each of these latents, effectively controlling how much that latent's direction contributes to the final output. -
Resulting Capability: The system can now explicitly control access to its parametric knowledge. During inference, it can dynamically suppress or amplify these specific directions based on the context, directly mitigating hallucinations caused by fine-tuning on unknown facts.
- Implement a
Monotonic Drift
Monitor for Post-Training:
-
Implementation: Integrate the MoRFI algorithm (Algorithm 1) as a standard evaluation step during any fine-tuning or RLHF process. After each epoch, compute the activation tensor for a held-out set of
known
facts. Run MoRFI to identify latents that are monotonically increasing or decreasing in response to the training data. If the drift of these latents exceeds a predefined threshold, pause training and issue a warning. -
Resulting Capability: The system can now proactively detect when it is starting to
forget
or lose access to its pre-trained knowledge, even if overall task accuracy is still improving. This prevents the silent degradation of general knowledge during specialized fine-tuning.
- Create a
Knowledge Recovery
Inference-Time Intervention:
-
Implementation: Build a post-processing module that can be applied to any fine-tuned model. This module uses the paper's findings to identify the optimal steering direction (positive/negative) and strength (α) for the top-10 most impactful latents (e.g., from Table 10). When a user query is identified as a closed-book QA task, the module applies the optimal steering intervention to the residual stream at the middle layer before the final forward pass.
-
Resulting Capability: The system can recover up to 85% of the knowledge that was
lost
during fine-tuning, without any retraining. This provides a powerful, inference-time tool to fix hallucination issues in already-deployed models.
- Develop a
Knowledge Access
Audit Tool:
-
Implementation: Create a diagnostic tool that, for a given model and a set of questions, uses the SAE to decompose the residual stream at the middle layer. The tool then compares the activation pattern against the known
increasing
anddecreasing
latent sets identified by MoRFI. A high activation ofincreasing
latents (associated with unknown facts) and low activation ofdecreasing
latents (associated with known facts) indicates a high risk of hallucination. -
Resulting Capability: The system can now provide a confidence score for its own answers based on internal mechanistic states, not just token probabilities. This allows for more reliable abstention or
I don't know
responses when the model's knowledge access pathways are disrupted.
- Add a
Fine-Tuning Data Purity
Checker:
-
Implementation: Before training on a new dataset, use the base model to compute
PCorrectfor all samples, classifying them as Known or Unknown. Then, use the paper's findings to predict the impact of the data mixture. If the percentage of Unknown facts exceeds a threshold (e.g., 25%), the system will flag the dataset as high-risk for inducing hallucinations and suggest either filtering out the Unknown samples or using a different training strategy. -
Resulting Capability: The system can now prevent hallucinations at the source by ensuring that fine-tuning data is aligned with the model's existing parametric knowledge, or by warning the user of the potential consequences before training begins.
What the Improved AI System Can Do:
-
Self-Diagnose Hallucinations: It can identify when its internal knowledge access pathways are compromised, rather than just when its output is likely wrong.
-
Recover Lost Knowledge: It can use targeted, single-latent interventions to restore access to pre-trained knowledge that was suppressed during fine-tuning, achieving significant accuracy gains.
-
Provide Mechanistic Confidence Scores: It can generate confidence scores based on the state of its internal knowledge representations, offering a more robust alternative to softmax probabilities.
-
Prevent Forgetting: It can monitor its own internal state during training and alert developers to the onset of knowledge degradation, allowing for early intervention.
-
Audit Its Own Training Data: It can assess the risk of a new dataset before training, preventing the introduction of hallucination-inducing knowledge conflicts.
Abstract
Large language models (LLMs) acquire most of their factual knowledge during the pre-training stage, through next token prediction. Subsequent stages of post-training often introduce new facts outwith the parametric knowledge, giving rise to hallucinations. While it has been demonstrated that supervised fine-tuning (SFT) on new knowledge may exacerbate the problem, the underlying mechanisms are still poorly understood. We conduct a controlled fine-tuning experiment, focusing on closed-book QA, and find latent directions that causally contribute to hallucinations. Specifically, we fine-tune Llama 3.1 8B, Gemma 2 9B and Mistral 7B v03 on seven distinct single QA datasets, controlling for the percentage of new knowledge and number of training epochs. By measuring performance on the test set, we validate that incrementally introducing new knowledge increases hallucinations, with the effect being more pronounced with prolonged training. We leverage pre-trained sparse autoencoders (SAEs) to analyze residual stream activations across various checkpoints for each model and propose Monotonic Relationship Feature Identification (MoRFI) for capturing causally relevant latents. MoRFI filters SAE features that respond monotonically to controlled fine-tuning data mixtures of a target property. Our findings show that exposure to unknown facts disrupts the model's ability to retrieve stored knowledge along a set of directions in the residual stream. Our pipeline reliably discovers them across distinct models, recovering knowledge through single-latent interventions.
Sources
- PaLM 2 Technical Report
- Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs
- Sparse Autoencoders Find Highly Interpretable Features in Language Models
- Not All Language Model Features Are One-Dimensionally Linear
- Do I Know This Entity? Knowledge Awareness and Hallucinations in Language Models
- Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?
- The Llama 3 Herd of Models
- The False Promise of Imitating Proprietary LLMs
- Llama Scope: Extracting Millions of Features from Llama-3.1-8B with Sparse Autoencoders
- A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions
- Mistral 7B
- Language Models (Mostly) Know What They Know
- Unfamiliar Finetuning Examples Control How Language Models Hallucinate
- The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets
- Narrow Finetuning Leaves Clearly Readable Traces in Activation Differences
- Language Models as Knowledge Bases?
- Simple Entity-Centric Questions Challenge Dense Retrievers
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Gemma 2: Improving Open Language Models at a Practical Size
- Model Organisms for Emergent Misalignment
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering