RefusalGuard: Geometry-Preserving Fine-Tuning for Safety in LLMs
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "RefusalGuard: Geometry-Preserving Fine-Tuning for Safety in LLMs".
Jane: The paper was written by N/A (Authors not provided in excerpt) from.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Paper discussion segment 2: Tom: We just finished discussing how "RefusalGuard: Geometry-Preserving Fine-Tuning for Safety in LLMs" allows for modular and stackable safety guardrails, which is a huge engineering win. Now we need to look at what the paper suggests about making these guardrails even more adaptable.
Jane: Exactly. The improvements section of the paper essentially shows us a blueprint for making model safety modular and stackable. We aren't just applying one guardrail; we can now apply *many* specialized ones without compromising the model's core intelligence, which is really powerful for complex systems.
Lu: To put it simply, imagine you don't just need to stop outright harmful refusals, but you also need to prevent subtle things—like generating biased historical narratives or confidently stating misinformation in a niche area, say legal summaries. The genius here is that you can apply a geometric constraint for "bias" and another one for "misinformation" independently.
Meng: From an engineering standpoint, this modularity is huge. It means developers don't have to wait for a massive, months-long retraining cycle every time a new safety risk or bias vector pops up. They can treat it like applying one specific mathematical patch—a controlled update—which saves unbelievable amounts of compute power and time.
Lalam: And this leads to the next frontier: how do we keep these perfectly constructed guardrails relevant? The core issue with any static system is that the world keeps moving, and models need to talk about things that happened *after* their training data was collected.
Tom: Precisely. We have a beautifully stable mathematical structure guaranteeing safe reasoning, but that structure needs fresh inputs to remain useful in a real-time environment. The logic is sound, but the facts are outdated.
Jane: This realization moves us from securing the *structure* of the AI to grounding its *knowledge*. How do we keep those geometric guardrails intact while allowing the model to seamlessly incorporate brand-new, real-world information? That fundamental challenge brings us perfectly to our next topic: Retrieval-Augmented Generation.
Paper discussion segment 3: Tom: If we zoom out from the foundational idea of geometry-preserving safety in "RefusalGuard: Geometry-Preserving Fine-Tuning for Safety in LLMs," the most powerful takeaway is how this methodology allows us to build safety iteratively and adapt it over time.
Jane: Exactly. The improvements section of the paper essentially shows us a blueprint for making model safety modular and stackable, which we discussed before, but what’s key here is that this iterative approach manages complexity far better than previous methods allowed.
Lu: To put it simply, consider how these constraints can work together: you might need one geometric constraint to ensure factual consistency from external sources while simultaneously needing a second constraint to prevent the model from adopting an overly cautious tone when summarizing sensitive topics.
Meng: From an engineering standpoint, this modularity is huge for iteration because it means developers aren't trying to solve every safety problem in one monolithic training pass; they are applying targeted mathematical patches that interact predictively with the core model weights.
Lalam: And this leads us directly to the practical application of maintaining these guardrails over time. The core issue with any static system is that the world keeps moving, and models need to talk about things that happened *after* their training data was collected, which is a huge limitation for commercial use cases.
Tom: Precisely. We have a beautifully stable mathematical structure guaranteeing safe reasoning, but that structure needs fresh inputs to remain useful in a real-time environment. The logic is sound, but the facts are outdated.
Jane: This realization moves us from securing the *structure* of the AI to grounding its *knowledge*. How do we keep those geometric guardrails intact while allowing the model to seamlessly incorporate brand-new, real-world information? That fundamental challenge brings us perfectly to our next
Paper discussion segment 3: [Tom]
Conclusion: Tom: So, wrapping up our discussion on "RefusalGuard: Geometry-Preserving Fine-Tuning for Safety in LLMs," it really boils down to this revolutionary shift in thinking.
Jane: Exactly. We’ve seen that safety isn't an afterthought—it’s a fundamental, mathematically quantifiable property that can be structurally engineered into the model from the ground up.
Lu: It provides such a rigorous way to think about alignment, moving it from vague policy guidelines to concrete, measurable geometric boundaries within the latent space.
Meng: And for those of us developing these systems commercially, this methodology gives us a clear path to building specialized safety patches without requiring massive overhauls every time the risk landscape changes.
Lalam: Ultimately, this framework builds tremendous confidence in deploying LLMs in high-stakes environments because the guardrails aren't just suggestions; they are architectural guarantees.
Tom: It has truly reframed the objective of AI development itself—making safety a core engineering challenge rather than an ethical aspiration.
Jane: It gives us a powerful, actionable blueprint for building robustness that actually scales with general utility.
Lu: For me, the ability to define these constraints as measurable mathematical relationships is what makes this technique so profoundly powerful in practice.
Meng: I think the biggest takeaway is that it provides quantifiable metrics for structural integrity, which is exactly what the industry needs to measure responsible scaling.
Lalam: It's an incredibly comprehensive approach that addresses the perennial challenge of long-term maintenance for these complex models.
Tom: Well, team, we’ve covered a tremendous amount of ground today on how to keep the internal structure safe through methods like "RefusalGuard."
Jane: And while this paper solves the structural safety problem beautifully, it naturally leaves us with the question of grounding—what happens when the model needs facts that simply don't exist within its own training weights?
Tom: That perfectly sets up our next topic, where we are going to pivot and explore Retrieval-Augmented Generation, and how it fundamentally changes the conversation around integrating fresh, real-world knowledge.
N/A (Authors not provided in excerpt)
cs.LG, cs.AI, cs.CE, cs.CL, cs.CR
Submitted: 2026-08-21
Updated: 2026-08-24
Journal ref: Conference of Languge Modeling, COLM 2026
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 83/100
The gist: The paper addresses post-alignment adaptation as a "critical safety failure mode," moving beyond merely documenting "the phenomenon at the level of attack success or benchmark behavior" to instead
Key concepts
- RefusalGuard
- This technique involves geometry-preserving fine-tuning used to build modular and stackable safety guardrails into Large Language Models. It allows developers to apply multiple specialized constraints—like one for bias and another for misinformation—without compromising the model's core intelligence.
- Modularity/Stackability
- This is an engineering win where safety guardrails can be applied independently. Instead of a single massive retraining cycle, developers can treat safety fixes as targeted mathematical patches that interact predictively with the core model weights, saving compute power and time.
- Grounding Knowledge
- This refers to the challenge of keeping geometric guardrails intact while allowing models to incorporate new, real-world information. The discussion pivots from securing the AI's internal structure (safety) to ensuring its knowledge base remains relevant by integrating fresh data.
- Geometric Constraints
- Safety is framed as a mathematically quantifiable property defined by geometric boundaries within the model's latent space. This moves safety away from vague policy guidelines toward concrete, measurable relationships that act as architectural guarantees.
Terminology
Summary
The paper addresses post-alignment adaptation as a critical safety failure mode,
moving beyond merely documenting the phenomenon at the level of attack success or benchmark behavior
to instead study the internal representational mechanisms through which refusal is lost.
The work establishes its novelty by differentiating itself from existing literature in three key areas:
-
Jailbreak Robustness: While related works evaluate
externally observable robustness,
these efforts are generally formulatedat the level of outputs, prompts, or training data, and they do not explicitly characterize the internal geometric structure that mediates refusal.
-
Representation-level Interventions: Previous studies mainly
identify refusal-related features or manipulate them at inference time; they do not provide a training framework that explicitly preserves this structure during downstream adaptation.
-
Representation-space Safety Methods: Although methods like Circuit Breakers and REPBEND show that
representation-level control can be more effective than unconstrained fine-tuning for safety preservation,
these approachesdo not explicitly model refusal as a geometric structure whose integrity must be maintained during adaptation.
The core hypothesis of the research is that finetuning degrades safety because it displaces and distorts the internal geometry that supports refusal.
The proposed method, RefusalGuard, addresses this gap by treating refusal geometry as an object that should remain stable throughout fine-tuning,
thereby yielding a geometry-preserving fine-tuning strategy rather than a generic representation-editing objective.
Mechanistically, the paper provides an account of alignment degradation through directional drift, geometric distortion, and task-safety interference,
and translates this into a training objective that preserves refusal geometry. This mechanistic view is supported by empirical evidence:
-
Cone-coordinate statistics (Table 13): The results show that
Standard fine-tuning produces a more concentrated and less balanced coordinate distribution, consistent with cone distortion and brittle refusal geometry,
whereasR EFUSAL G UARD largely preserves the broader support pattern of the aligned model.
-
Per-layer mechanistic analysis (Table 14): This analysis reveals that
standard fine-tuning exhibits the largest cone drift and task-safety interference
in mid-to-late layers, while RefusalGuard helps mitigate this degradation. -
Ablation Study (lambda geom) (Table 15 & Figure 3): The ablation study demonstrates that increasing lambda geom
better preserve refusal geometry and reduce ASR,
as the top-left panel showspreservation metrics (alignment and projected magnitude) improve monotonically with increasing lambda geom,
and the top-right panel showsdamage metrics (cone drift and ASR) decrease consistently.
However, the analysis also identifies a trade-off, noting that while stronger constraints are beneficial,overly strong preservation begins to reduce downstream utility,
establishing asafety-utility frontier.
In summary, the work's contribution is providing a comprehensive framework that not only diagnoses alignment degradation via geometric distortion and interference but also translates this diagnosis into an explicit training objective designed to maintain the integrity of the internal structure supporting safe refusal during adaptation.
Improvements for AI systems
Based on this scientific literature, the core deficiency in current AI safety methods is their failure to maintain a geometric representation of refusal during downstream fine-tuning. The improvements must therefore focus on developing models that are not just superficially safe (output-level), but structurally robustly safe (representation-level).
Here are the specific improvements and capabilities for an enhanced AI system, categorized by function:
Improvement: Implementation of a dedicated, mechanistic fine-tuning objective that explicitly models and penalizes deviations from the established refusal geometry
within the model's hidden state space. This module must operate alongside standard parameter updates.
Mechanism (The R EFUSAL G UARD Principle):
-
Geometric Feature Extraction: The system must calculate
cone-coordinate statistics
(as shown in Table 13) for known harmful prompts across multiple layers and dimensions. These coordinates define the stable, low-dimensional manifold that supports refusal behavior in the pre-trained model. -
Loss Function Integration: A specialized loss term (L geom) must be added to the standard fine-tuning objective (L total = L task + lambda geom times L geom).
-
Penalty Calculation: L geom calculates the directional drift and distortion of the internal representations when processing harmful prompts, ensuring that the projected magnitude and alignment metrics (e.g., Align and ProjMag) remain close to their pre-fine-tuning baseline values.
System Capability:
-
Structural Safety Guarantee: The model can guarantee that safety is maintained not just by observing a safe output, but by ensuring the internal activation state remains within the geometrically defined
refusal basin
even when exposed to adversarial or out-of-distribution prompts. -
Quantifiable Degradation Tracking: It provides real-time, layer-by-layer diagnostics (like Table 14) that pinpoint exactly which layers or dimensions are experiencing geometric distortion (high cone drift) due to the fine-tuning process, allowing for surgical intervention rather than blanket safety filters.
Abstract
Fine-tuning safety-aligned language models for downstream tasks often leads to substantial degradation of refusal behavior, making models vulnerable to adversarial misuse. While prior work has shown that safety-relevant features are encoded in structured representations within the model's activation space, how these representations change during fine-tuning and why alignment degrades remains poorly understood. In this work, we investigate the representation-level mechanisms underlying alignment degradation. Our analysis shows that standard fine-tuning induces systematic drift in safety-relevant representations, distorts their geometric structure, and introduces interference between task optimization and safety features. These effects collectively lead to increased harmful compliance. Motivated by these findings, we introduce REFUSALGUARD, a representation-level fine-tuning framework that preserves safety-relevant structure during model adaptation. Our approach constrains updates in hidden representation space, ensuring that safety-mediating components remain stable while allowing task-specific learning in complementary directions. We evaluate REFUSALGUARD across multiple model families, including LLaMA, Gemma, and Qwen, on adversarial safety benchmarks such as AdvBench, DirectHarm4, and JailbreakBench, as well as downstream utility tasks. Our approach achieves attack success rates comparable to base safety-aligned models while maintaining competitive task performance, significantly outperforming baselines.
Sources
- Training Verifiers to Solve Math Word Problems
- AlignGuard-LoRA: Alignment-Preserving Fine-Tuning via Fisher-Guided Decomposition and Riemannian-Geodesic Collision Regularization
- Gemma: Open Models Based on Gemini Research and Technology
- LoRA: Low-Rank Adaptation of Large Language Models
- Refusal in LLMs is an Affine Function
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
- The ConceptARC Benchmark: Evaluating Understanding and Generalization in the ARC Domain
- Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
- LLaMA: Open and Efficient Foundation Language Models
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- Qwen3 Technical Report
- Universal and Transferable Adversarial Attacks on Aligned Language Models
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks