RefusalGuard: Geometry-Preserving Fine-Tuning for Safety in LLMs

arXiv:2605.01913 · cs.LG, cs.AI, cs.CE, cs.CL, cs.CR · Submitted 2026-08-21 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "RefusalGuard: Geometry-Preserving Fine-Tuning for Safety in LLMs".

Jane: The paper was written by N/A (Authors not provided in excerpt) from.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Paper discussion segment 2: Tom: We just finished discussing how "RefusalGuard: Geometry-Preserving Fine-Tuning for Safety in LLMs" allows for modular and stackable safety guardrails, which is a huge engineering win. Now we need to look at what the paper suggests about making these guardrails even more adaptable.

Jane: Exactly. The improvements section of the paper essentially shows us a blueprint for making model safety modular and stackable. We aren't just applying one guardrail; we can now apply *many* specialized ones without compromising the model's core intelligence, which is really powerful for complex systems.

Lu: To put it simply, imagine you don't just need to stop outright harmful refusals, but you also need to prevent subtle things—like generating biased historical narratives or confidently stating misinformation in a niche area, say legal summaries. The genius here is that you can apply a geometric constraint for "bias" and another one for "misinformation" independently.

Meng: From an engineering standpoint, this modularity is huge. It means developers don't have to wait for a massive, months-long retraining cycle every time a new safety risk or bias vector pops up. They can treat it like applying one specific mathematical patch—a controlled update—which saves unbelievable amounts of compute power and time.

Lalam: And this leads to the next frontier: how do we keep these perfectly constructed guardrails relevant? The core issue with any static system is that the world keeps moving, and models need to talk about things that happened *after* their training data was collected.

Tom: Precisely. We have a beautifully stable mathematical structure guaranteeing safe reasoning, but that structure needs fresh inputs to remain useful in a real-time environment. The logic is sound, but the facts are outdated.

Jane: This realization moves us from securing the *structure* of the AI to grounding its *knowledge*. How do we keep those geometric guardrails intact while allowing the model to seamlessly incorporate brand-new, real-world information? That fundamental challenge brings us perfectly to our next topic: Retrieval-Augmented Generation.

Paper discussion segment 3: Tom: If we zoom out from the foundational idea of geometry-preserving safety in "RefusalGuard: Geometry-Preserving Fine-Tuning for Safety in LLMs," the most powerful takeaway is how this methodology allows us to build safety iteratively and adapt it over time.

Jane: Exactly. The improvements section of the paper essentially shows us a blueprint for making model safety modular and stackable, which we discussed before, but what’s key here is that this iterative approach manages complexity far better than previous methods allowed.

Lu: To put it simply, consider how these constraints can work together: you might need one geometric constraint to ensure factual consistency from external sources while simultaneously needing a second constraint to prevent the model from adopting an overly cautious tone when summarizing sensitive topics.

Meng: From an engineering standpoint, this modularity is huge for iteration because it means developers aren't trying to solve every safety problem in one monolithic training pass; they are applying targeted mathematical patches that interact predictively with the core model weights.

Lalam: And this leads us directly to the practical application of maintaining these guardrails over time. The core issue with any static system is that the world keeps moving, and models need to talk about things that happened *after* their training data was collected, which is a huge limitation for commercial use cases.

Tom: Precisely. We have a beautifully stable mathematical structure guaranteeing safe reasoning, but that structure needs fresh inputs to remain useful in a real-time environment. The logic is sound, but the facts are outdated.

Jane: This realization moves us from securing the *structure* of the AI to grounding its *knowledge*. How do we keep those geometric guardrails intact while allowing the model to seamlessly incorporate brand-new, real-world information? That fundamental challenge brings us perfectly to our next

Paper discussion segment 3: [Tom]

Conclusion: Tom: So, wrapping up our discussion on "RefusalGuard: Geometry-Preserving Fine-Tuning for Safety in LLMs," it really boils down to this revolutionary shift in thinking.

Jane: Exactly. We’ve seen that safety isn't an afterthought—it’s a fundamental, mathematically quantifiable property that can be structurally engineered into the model from the ground up.

Lu: It provides such a rigorous way to think about alignment, moving it from vague policy guidelines to concrete, measurable geometric boundaries within the latent space.

Meng: And for those of us developing these systems commercially, this methodology gives us a clear path to building specialized safety patches without requiring massive overhauls every time the risk landscape changes.

Lalam: Ultimately, this framework builds tremendous confidence in deploying LLMs in high-stakes environments because the guardrails aren't just suggestions; they are architectural guarantees.

Tom: It has truly reframed the objective of AI development itself—making safety a core engineering challenge rather than an ethical aspiration.

Jane: It gives us a powerful, actionable blueprint for building robustness that actually scales with general utility.

Lu: For me, the ability to define these constraints as measurable mathematical relationships is what makes this technique so profoundly powerful in practice.

Meng: I think the biggest takeaway is that it provides quantifiable metrics for structural integrity, which is exactly what the industry needs to measure responsible scaling.

Lalam: It's an incredibly comprehensive approach that addresses the perennial challenge of long-term maintenance for these complex models.

Tom: Well, team, we’ve covered a tremendous amount of ground today on how to keep the internal structure safe through methods like "RefusalGuard."

Jane: And while this paper solves the structural safety problem beautifully, it naturally leaves us with the question of grounding—what happens when the model needs facts that simply don't exist within its own training weights?

Tom: That perfectly sets up our next topic, where we are going to pivot and explore Retrieval-Augmented Generation, and how it fundamentally changes the conversation around integrating fresh, real-world knowledge.

N/A (Authors not provided in excerpt)

cs.LG, cs.AI, cs.CE, cs.CL, cs.CR

Submitted: 2026-08-21

Updated: 2026-08-24

Journal ref: Conference of Languge Modeling, COLM 2026

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 83/100

The gist: The paper addresses post-alignment adaptation as a "critical safety failure mode," moving beyond merely documenting "the phenomenon at the level of attack success or benchmark behavior" to instead

Key concepts

RefusalGuard
This technique involves geometry-preserving fine-tuning used to build modular and stackable safety guardrails into Large Language Models. It allows developers to apply multiple specialized constraints—like one for bias and another for misinformation—without compromising the model's core intelligence.
Modularity/Stackability
This is an engineering win where safety guardrails can be applied independently. Instead of a single massive retraining cycle, developers can treat safety fixes as targeted mathematical patches that interact predictively with the core model weights, saving compute power and time.
Grounding Knowledge
This refers to the challenge of keeping geometric guardrails intact while allowing models to incorporate new, real-world information. The discussion pivots from securing the AI's internal structure (safety) to ensuring its knowledge base remains relevant by integrating fresh data.
Geometric Constraints
Safety is framed as a mathematically quantifiable property defined by geometric boundaries within the model's latent space. This moves safety away from vague policy guidelines toward concrete, measurable relationships that act as architectural guarantees.

Terminology

Summary

The paper addresses post-alignment adaptation as a critical safety failure mode, moving beyond merely documenting the phenomenon at the level of attack success or benchmark behavior to instead study the internal representational mechanisms through which refusal is lost.

The work establishes its novelty by differentiating itself from existing literature in three key areas:

  1. Jailbreak Robustness: While related works evaluate externally observable robustness, these efforts are generally formulated at the level of outputs, prompts, or training data, and they do not explicitly characterize the internal geometric structure that mediates refusal.

  2. Representation-level Interventions: Previous studies mainly identify refusal-related features or manipulate them at inference time; they do not provide a training framework that explicitly preserves this structure during downstream adaptation.

  3. Representation-space Safety Methods: Although methods like Circuit Breakers and REPBEND show that representation-level control can be more effective than unconstrained fine-tuning for safety preservation, these approaches do not explicitly model refusal as a geometric structure whose integrity must be maintained during adaptation.

The core hypothesis of the research is that finetuning degrades safety because it displaces and distorts the internal geometry that supports refusal. The proposed method, RefusalGuard, addresses this gap by treating refusal geometry as an object that should remain stable throughout fine-tuning, thereby yielding a geometry-preserving fine-tuning strategy rather than a generic representation-editing objective.

Mechanistically, the paper provides an account of alignment degradation through directional drift, geometric distortion, and task-safety interference, and translates this into a training objective that preserves refusal geometry. This mechanistic view is supported by empirical evidence:

  • Cone-coordinate statistics (Table 13): The results show that Standard fine-tuning produces a more concentrated and less balanced coordinate distribution, consistent with cone distortion and brittle refusal geometry, whereas R EFUSAL G UARD largely preserves the broader support pattern of the aligned model.

  • Per-layer mechanistic analysis (Table 14): This analysis reveals that standard fine-tuning exhibits the largest cone drift and task-safety interference in mid-to-late layers, while RefusalGuard helps mitigate this degradation.

  • Ablation Study (lambda geom) (Table 15 & Figure 3): The ablation study demonstrates that increasing lambda geom better preserve refusal geometry and reduce ASR, as the top-left panel shows preservation metrics (alignment and projected magnitude) improve monotonically with increasing lambda geom, and the top-right panel shows damage metrics (cone drift and ASR) decrease consistently. However, the analysis also identifies a trade-off, noting that while stronger constraints are beneficial, overly strong preservation begins to reduce downstream utility, establishing a safety-utility frontier.

In summary, the work's contribution is providing a comprehensive framework that not only diagnoses alignment degradation via geometric distortion and interference but also translates this diagnosis into an explicit training objective designed to maintain the integrity of the internal structure supporting safe refusal during adaptation.

Improvements for AI systems

Based on this scientific literature, the core deficiency in current AI safety methods is their failure to maintain a geometric representation of refusal during downstream fine-tuning. The improvements must therefore focus on developing models that are not just superficially safe (output-level), but structurally robustly safe (representation-level).

Here are the specific improvements and capabilities for an enhanced AI system, categorized by function:


Improvement: Implementation of a dedicated, mechanistic fine-tuning objective that explicitly models and penalizes deviations from the established refusal geometry within the model's hidden state space. This module must operate alongside standard parameter updates.

Mechanism (The R EFUSAL G UARD Principle):

  1. Geometric Feature Extraction: The system must calculate cone-coordinate statistics (as shown in Table 13) for known harmful prompts across multiple layers and dimensions. These coordinates define the stable, low-dimensional manifold that supports refusal behavior in the pre-trained model.

  2. Loss Function Integration: A specialized loss term (L geom) must be added to the standard fine-tuning objective (L total = L task + lambda geom times L geom).

  3. Penalty Calculation: L geom calculates the directional drift and distortion of the internal representations when processing harmful prompts, ensuring that the projected magnitude and alignment metrics (e.g., Align and ProjMag) remain close to their pre-fine-tuning baseline values.

System Capability:

  • Structural Safety Guarantee: The model can guarantee that safety is maintained not just by observing a safe output, but by ensuring the internal activation state remains within the geometrically defined refusal basin even when exposed to adversarial or out-of-distribution prompts.

  • Quantifiable Degradation Tracking: It provides real-time, layer-by-layer diagnostics (like Table 14) that pinpoint exactly which layers or dimensions are experiencing geometric distortion (high cone drift) due to the fine-tuning process, allowing for surgical intervention rather than blanket safety filters.

Abstract

Fine-tuning safety-aligned language models for downstream tasks often leads to substantial degradation of refusal behavior, making models vulnerable to adversarial misuse. While prior work has shown that safety-relevant features are encoded in structured representations within the model's activation space, how these representations change during fine-tuning and why alignment degrades remains poorly understood. In this work, we investigate the representation-level mechanisms underlying alignment degradation. Our analysis shows that standard fine-tuning induces systematic drift in safety-relevant representations, distorts their geometric structure, and introduces interference between task optimization and safety features. These effects collectively lead to increased harmful compliance. Motivated by these findings, we introduce REFUSALGUARD, a representation-level fine-tuning framework that preserves safety-relevant structure during model adaptation. Our approach constrains updates in hidden representation space, ensuring that safety-mediating components remain stable while allowing task-specific learning in complementary directions. We evaluate REFUSALGUARD across multiple model families, including LLaMA, Gemma, and Qwen, on adversarial safety benchmarks such as AdvBench, DirectHarm4, and JailbreakBench, as well as downstream utility tasks. Our approach achieves attack success rates comparable to base safety-aligned models while maintaining competitive task performance, significantly outperforming baselines.

Sources

Related papers