Obliviate: Erasing Concepts from Autoregressive Image Generation Models

arXiv:2606.28643 · cs.CV · Submitted 2026-06-26 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Obliviate: Erasing Concepts from Autoregressive Image Generation Models".

Jane: Obliviate introduces a guidance-based concept erasure method for autoregressive image generation models, addressing the safety gap in these architectures by removing targeted concepts like nudity, gory content,

Tom: First, who's behind it and why it matters.

Title and authors: Tom: So, we're talking about the paper 'Obliviate: Erasing Concepts from Autoregressive Image Generation Models', and honestly, the title tells you exactly what this research is about—it's all about taking generative AI models that create images sequentially, like autoregressive models, and teaching them how to reliably remove specific concepts from those images. Jane, what do you make of that?

Jane: It sounds like they are tackling a real problem in the area because these image generators are getting so good at making realistic pictures that we need ways to keep harmful or inappropriate content out of circulation. The authors, Hossein Shakibania and his team, seem to be focusing on filling a gap where concept erasure for autoregressive generation hasn't been explored much before.

Lu: I think the core idea they are pushing is moving concept erasure from diffusion models over to these sequential generation models, which is a significant hurdle because those two architectures operate in fundamentally different ways. It's about figuring out how to impose that kind of control onto a process that builds an image one token at a time.

Meng: From an engineering standpoint, I'm interested in the specifics of their approach; if they can do this effectively on autoregressive models, it means we don't have to rely solely on external filters or post-processing steps to catch things like nudity or branding.

Lalam: I think this work is important because if we can control what these generative models produce internally through training, it really helps build a safer and more trustworthy culture for how AI is used in creating visual content.

The paper's summary: Tom: So, to summarize what the Obliviate paper actually does, they introduce a guidance-based concept erasure method specifically designed for autoregressive image generation models. They outline three main design choices they made to bridge the gap between diffusion methods and autoregressive settings.

Jane: That sounds like a lot of technical steps, but I think it boils down to aligning the training targets better by using KL-based supervision over visual token distributions, doing trajectory-level updates across full rollouts instead of just single tokens, and using aligned visual prefixes to help stabilize the target construction.

Lu: Exactly. They start by defining a teacher target based on contrasting conditional and unconditional next-token predictions at each position, which is similar to how it's done in diffusion models, but they adapt it for this setting. The key is creating that pseudo-unconditional branch that shares the same visual prefix as the main conditional branch, which helps isolate the tokens responsible for sustaining the undesired concept.

Meng: So they are essentially forcing the model to learn a distribution match between what it *should* generate and what represents a safe outcome across an entire generation path. That sounds computationally intensive to train, though.

Lalam: It’s about shifting from looking at isolated mistakes in one step to supervising the whole sequence of tokens, which I think is much more effective for catching complex patterns like full scenes of inappropriate content or persistent branding elements.

The paper's improvements: Tom: Moving on to what they actually improved, the authors detail a few key methodological enhancements that make Obliviate work where previous methods struggled. They focused on moving from single-step objectives to trajectory-wide supervision and using Kullback–Leibler divergence over predicted distributions instead of just hard token labels.

Jane: That KL divergence loss is interesting because it doesn't force the model to pick one specific safe token, but rather encourages the student model to match the teacher's target distribution over the full sampled rollout via a KL divergence over visual logits. It’s like telling it, "match this entire range of possibilities rather than just picking one safe word."

Lu: And they also show how to overcome those obstacles by using trajectory-level updates over full rollouts, which is motivated by the idea that harmful behavior unfolds across the generation chain rather than at an isolated token. This aligns with prior work showing better erasure quality from updates along complete generation paths.

Meng: So, they're essentially training the student model to suppress harmful continuations by matching a target distribution over many steps, which suggests a more nuanced way to handle safety constraints than just filtering specific tokens out.

Lalam: It's powerful because this method allows for distribution-level supervision, meaning it suppresses harmful continuations while still preserving the safe semantics of the overall image generation process.

Conclusion: Tom: So, wrapping up what we've heard on 'Obliviate: Erasing Concepts from Autoregressive Image Generation Models', it seems this guidance-based concept erasure method has achieved strong empirical results across several scenarios, successfully reducing concept detection rates on benchmarks like T2I-RiskyPrompt to around three point one five when using the Liquid model <ref:2606.28643#pg0,Obliviate: Erasing Concepts from Autoregressive Image Generation Models>.

Jane: It really shows that by focusing on aligning visual prefixes and using distribution matching over full rollouts, we can make significant improvements in removing things like explicit content and branded imagery without sacrificing the overall quality of the image generation.

Lu: The authors demonstrate that their approach works across three state-of-the-art autoregressive image generators, showing consistent performance in reducing detection rates for nudity, gory content, and brand symbols to near zero on those models.

Meng: From an engineering perspective, this means we have a new toolset for fine-tuning these models directly to enforce complex safety policies during the generation process itself rather than just cleaning up the output afterwards.

Lalam: I think the implication here is that we are moving towards generative AI systems that are inherently safer because the safety constraints are built into their core generation mechanism, which could drastically improve how we build and deploy these tools responsibly.

Tom: That’s a solid summary of what 'Obliviate: Erasing Concepts from Autoregressive Image Generation Models' has delivered. Thanks for tuning in to this deep dive into the research on content erasure.

Hossein Shakibania, Jonas Henry Grebe, Tobias Braun, Ege Aktemur, Saleh Aslani, Mehmet G. Yiğit, Marcus Rohrbach

TU Darmstadt

cs.CV

Submitted: 2026-06-26

Updated: 2026-10-04

Importance score: 90/100

The gist: Obliviate introduces a guidance-based concept erasure method for autoregressive image generation models, addressing the safety gap in these architectures by removing targeted concepts like nudity,

Key concepts

KL-based supervision
This technique uses Kullback–Leibler divergence to measure how much the model's predicted token distributions differ from a desired target distribution. It helps train the model to generate outputs that match a specific, safe pattern by penalizing deviations in the probability of predicting certain visual tokens.
Trajectory-level updates
Instead of just looking at individual steps, Obliviate trains the model based on the entire sequence of generated tokens (the full rollout). This is because harmful content often develops across many steps, so supervising the whole path leads to better erasure than just fixing single mistakes.
Aligned visual prefixes
The method reuses previously generated harmful sequences as a starting point for an unconditional branch. This alignment stabilizes the target concept by isolating the specific tokens responsible for sustaining that undesired element, making it easier to suppress them later.

Terminology

Summary

Obliviate introduces a guidance-based concept erasure method for autoregressive image generation models, addressing the safety gap in these architectures by removing targeted concepts like nudity, gory content, and brands while preserving overall generation quality.

How it works

Obliviate builds upon three key design choices to overcome the challenges of translating diffusion-based erasure methods to the autoregressive setting:

  1. KL-based supervision over visual token distributions.

  2. Trajectory-level updates over full autoregressive rollouts.

  3. Aligned visual prefixes for stable target construction.

The method starts by defining a teacher target based on contrasting conditional and unconditional next-token predictions at each position, as in the diffusion model formulation:

(4) p(k) tgt = softmax z(k) tgt, x

Key Idea 1 involves aligning visual prefixes. The paper states that We reuse the generated harmful trajectory as a shared prefix for the unconditional branch, thereby forming a pseudo-unconditional branch without text conditioning but with an aligned visual prefix. This alignment stabilizes the target and isolates the tokens most responsible for sustaining the undesired concept.

How it works (Continued)

Key Idea 2 focuses on training over full rollouts. The authors extend the single-step objective to trajectory-wide supervision:

(8) L FT-CE = (1/N)sum k=1 to N L CE(k) = -(1/N)sum k=1 to N log p theta(x

, c). This approach is motivated by the fact that harmful behavior unfolds over the generation chain rather than at an isolated token, and aligns with prior erasure work showing improved erasure quality from updates along complete generation paths.

Key Idea 3 involves distribution-level supervision. Instead of hard token-level supervision, Obliviate uses a Kullback–Leibler divergence over the predicted distributions:

(9) L method = (1/N)sum k=1 to N D KL(p tgt(k) p theta(k)). This trains the student to match the teacher-induced target distributions over the full sampled rollout via a KL divergence over the visual logits, which suppresses harmful continuations while preserving safe semantics.

The modified Safe Latent Diffusion (SLD) framework, denoted as SLD∗, is adapted for autoregressive models by removing warm-up and momentum. Instead of predicting noise estimates, the model outputs logits over a discrete vocabulary for the next token. The safety adjustment term γ isolates the target concept relative to the unconditional baseline:

**(14) γ = μ **

(z theta(x<k,c) - z theta(x<k,φ))

Finally, standard classifier-free guidance is modified to incorporate this safety adjustment, yielding the final safe target logits for the next token:

**(15) z final = z theta(x<k,…,φ) + s g **

(z theta(x<k,p) - z theta(x<k,…,φ) - γ)

The method is implemented using LoRA fine-tuning with rank 32 and specific dropout settings across three state-of-the-art models: Liquid, Emu3-Gen, and Janus-Pro. The guidance scale η is varied based on the task; for instance, moderate values around η =2.0 to 3.0 provide the best balance for explicit content erasure on Liquid.

The evaluation uses diverse benchmarks including T2I-RiskyPrompt (T2IRP), Inappropriate Image Prompts (I2P), Ring-A-Bell (RAB), and the Unbranding benchmark, with metrics such as Concept Detection Rate (CDR) and image quality assessments like FID and CLIP Score. Obliviate consistently outperforms baselines across explicit content, gory content, and unbranding scenarios. For example, on Liquid for explicit content erasure, it reduces the CDR from 45.82 to 3.73 on T2I-RP and from 91.58 to 3.15 on RAB while maintaining utility metrics like FID at 15.41 and CLIP at 13.06, substantially outperforming Negative Prompting and SLD∗ methods.

The ablation study shows that the choice of concept prompt matters: "Branding: simple prompt based on the exact brand name performs substantially better than the more descriptive alternative on all models.

Improvements for AI systems

Based on the scientific paper Obliviate: Erasing Concepts from Autoregressive Image Generation Models, here are specific, actionable improvements for AI systems and a description of what those improved systems can achieve.


The core innovation is the introduction of Obliviate, a guidance-based concept erasure method specifically designed for autoregressive (AR) image generation models.

Here are the specific improvements and capabilities:

  1. Improving Safety in Autoregressive Text-to-Image Models via Concept Erasure

  2. Enhancing Robustness Against Misuse of Generative AI

  3. Achieving High-Fidelity Image Generation While Enforcing Content Restrictions

Specific Improvements:

  1. Improve Safety in Autoregressive Text-to-Image Models via Concept Erasure

  2. Enhancing Robustness Against Misuse of Generative AI

  3. Achieving High-Fidelity Image Generation While Enforcing Content Restrictions

Detailed Capabilities of the Improved System (Obliviate):

The Obliviate system can perform targeted concept erasure on AR image generation models (like Liquid, Emu3-Gen, and Janus-Pro) to suppress specific visual concepts while preserving the overall semantic quality and utility of the generated image. This is achieved through a three-pronged architectural approach:

  1. An aligned pseudo-unconditional branch that shares the same visual prefix as the conditional branch to stabilize target construction (Key Idea 1).

  2. Trajectory-level updates over full autoregressive rollouts instead of isolated token updates, leveraging causal masking for parallel supervision (Key Idea 2).

  3. Distribution-level supervision via a Kullback–Leibler (KL) divergence loss over visual logits, which allows the model to reallocate probability mass away from harmful continuations toward safer alternatives (Key Idea 3).

Specific Outcomes and Specific Applications:

The improved AI system, Obliviate, can achieve the following specific results across various safety scenarios:

  1. Suppress explicit content (nudity) with high precision: Obliviate reduces the Concept Detection Rate (CDR) on benchmarks like T2I-RiskyPrompt and Ring-A-Bell from 91.58 down to 3.15 on the Liquid model, significantly outperforming state-of-the-art baselines like Negative Prompting and Supervised Fine-Tuning (SFT).

  2. Eradicate gory content: It substantially reduces the CDR for violent and bloody visual concepts, lowering detection rates from over 94% to around 39% on the Liquid model while maintaining competitive image quality metrics (FID/CLIP).

  3. Remove branded imagery with near-zero detection: Obliviate is exceptionally effective at unbranding. It reduces the Concept Detection Rate for specific brands (like Coca-Cola) to near zero across all evaluated models, demonstrating that its KL-based distribution matching suppresses not just the exact brand token but also visually similar alternative patterns.

  4. Preserve General Utility: Crucially, Obliviate achieves these safety gains while maintaining high visual fidelity. For example, on Janus-Pro for explicit content erasure, it reduces the Concept Detection Rate to 1.05 compared to 11.58 for the baseline EAR method, and its FID remains comparable to the original model (e.g., 12.31 vs. 31.63).

  5. Provide Adaptive Guidance Control: The method allows for fine-grained control over erasure strength using a negative guidance scale parameter (η), enabling researchers to balance safety against utility trade-offs based on the specific concept being erased (e.g., using η=2 for explicit content and η=10 for brand removal).

In summary, Obliviate transforms autoregressive image generation from a vulnerable system into one capable of reliably enforcing complex safety and content policies through targeted, in-model weight updates.

Sources

Related papers