Forensic-Aware Continual Adaptation for Image Forgery Localization

arXiv:2609.38251 · cs.CR, cs.AI · Submitted 2026-09-29 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.

Nadia: Today's paper: "Forensic-Aware Continual Adaptation for Image Forgery Localization".

Elias: The rapid evolution of image manipulation techniques has raised pressing public security concerns,

Nadia: First, who's behind it and why it matters.

Paper summary: Nadia: So we’ve covered the high-level concept, and now I want to go into a bit more detail about what this specific paper, "Forensic-Aware Continual Adaptation for Image Forgery Localization," actually claims regarding its core thesis. The main point is that existing Image Forgery Localization methods lack the ability to adapt dynamically when new forgery techniques appear in real-time, which is a major hurdle in practical security workflows.

Elias: Precisely, Nadia. The authors argue that because data typically arrives sequentially in forensic scenarios, current IFL models are fundamentally ill-equipped for continual learning without modification; they overlook the need for the model to evolve alongside the incoming data stream.

Priya: What they claim is that they introduce a novel continual learning framework to solve this gap, specifically designed so that IFL models can progressively adapt to new forgery domains while simultaneously retaining the forensic knowledge they acquired from previous datasets. This is framed as a practical necessity for real-world forensic applications where data isn't static.

Nadia: That retention of knowledge is key, and they propose two specific mechanisms to handle adaptation and preservation: first, a Spatial Mixture-of-Forensic-Experts module for representation adaptation, and second, a Fisher-weighted LoRA Gradient surgery strategy for knowledge preservation.

Elias: The SMoFE module focuses on fine-grained spatial routing by adaptively combining local anomalies from different forensic experts, using four extractors to capture diverse traces like edge discontinuities and compression blockiness. It's about selecting the most relevant cues based on both semantic features and forensic trace features.

Priya: And then there’s FEGDP, which transforms those low-level forensic traces into structured localization evidence for SAM by decomposing forgery evidence into regional inconsistency and boundary evidence, which is then used to construct a dense prompt. This bridges the gap between raw forensic data and the localization output.

Nadia: That sounds like they are building a pipeline that adapts its internal representation using spatial routing, then structures that evidence for downstream tasks like SAM, all while managing the learning process through those surgical updates. The paper claims this leads to state-of-the-art performance across two distinct continual learning protocols.

Elias: And their experimental setup under Protocol one and Protocol two is what validates this claim; they show that the proposed framework successfully navigates the sequential task arrival challenges, which is a crucial part of proving its viability in practice.

Priya: I'm looking at the benchmark structure itself, specifically how they define Protocol one with four manipulation datasets: Classic, Defacto, FantasticReality, and TampCOCO, sorted by release dates to simulate real deployment scenarios. This gives us a concrete view of the sequential challenges they are testing against.

Nadia: And Protocol two adds another layer by covering cross-content learning involving natural images, document images, and scientific images. This breadth in testing shows the framework isn't just tuned for one type of forgery but is more versatile across different image domains.

Elias: The paper’s main contribution is therefore not just proposing a single new technique, but a comprehensive continual learning framework that integrates representation adaptation with knowledge preservation through these specific modules.

Priya: I think the data they present, showing state-of-the-art results in both pixel-level localization and image-level detection across these varied protocols, really substantiates the authors' argument about its practical utility.

Nadia: So, to summarize what we’ve discussed for this paper is that it proposes a continual learning framework designed to make IFL models robust against evolving forgery techniques by using spatial mixture-of-forensic-experts for adaptation and Fisher-weighted LoRA Gradient surgery for knowledge preservation.

Elias: It’s a framework built on managing the trade-off between plasticity and stability in sequential learning scenarios, which addresses the fundamental difficulty they identified in existing IFL methods.

Priya: It really shows that we can develop tools that are capable of handling the complexities of real-world sequential data evolution effectively.

Conclusion: Nadia: Wrapping up our discussion on "Forensic-Aware Continual Adaptation for Image Forgery Localization," the title itself perfectly captures the goal—it’s about making IFL models aware of forensics while ensuring they can continually adapt to new challenges without forgetting what they've already learned. The authors are proposing a method that addresses the dynamic nature of image manipulation threats.

Elias: That’s right, and their work centers on solving the problem of catastrophic forgetting in continual learning by using targeted surgery on LoRA gradients, which is a very specific mechanism for preserving old knowledge during adaptation. The implications are that we need to consider how these importance estimates influence the learned model's behavior.

Priya: From a research standpoint, this paper contributes a concrete framework for handling sequential data streams in image forensics, moving the field toward more robust and adaptable detection tools capable of handling diverse forgery types effectively.

Nadia: The real-world implications are that security systems could deploy detectors that stay current with emerging manipulation techniques without needing constant, massive retraining cycles every time a new forgery style emerges.

Elias: If this framework holds up under rigorous testing across protocols like the ones they benchmarked, it suggests a more stable foundation for developing forensic AI tools in dynamic environments.

Priya: It opens up avenues for further research into how these forensic evidence-guided prompting and adaptation mechanisms can be generalized to other sequential learning problems outside of just image forensics.

Nadia: That’s the big picture we’ve been discussing, showing a method that provides a solid structure for building next-generation, continuously learning security tools against evolving threats.

Chenqi Kong, Song Xia, Anwei Luo, Peisong He, Alex C. Kot, Yuming Fang

School of Computing, National University of Singapore · ROSE Lab, School of EEE, Nanyang Technological University

cs.CR, cs.AI

Submitted: 2026-09-29

Updated: 2026-09-29

Code: https://github.com/mjkwon2021/CAT-Net

Project page: https://defactodataset.github.io

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 91/100

The gist: The rapid evolution of image manipulation techniques has raised pressing public security concerns, and existing Image Forgery Localization (IFL) methods often fail to adapt dynamically to newly

Key concepts

Spatial Mixture-of-Forensic-Experts (SMoFE)
This module adapts the model's understanding of forensic traces by combining information from four different expert trace extractors. It uses a gate network to decide which specific local anomalies—like edge discontinuities or compression blockiness—are most relevant for a given image, creating a fused forensic embedding.
Fisher-Weighted LoRA Gradient (FLAG) Surgery
This strategy preserves old knowledge during adaptation by analyzing the importance of different learned parameters. It uses Fisher information to weight updates, selectively suppressing conflicting changes in dimensions related to previous tasks. This ensures that critical forensic knowledge from earlier datasets is not accidentally forgotten when learning new ones.
Forensic Evidence-Guided Dense Prompting (FEGDP)
This technique translates low-level forensic traces into structured localization evidence suitable for vision models like SAM. It separates evidence into regional inconsistency and boundary patterns, creating a detailed prompt that explicitly shows the model where the forgery is located and how it transitions.
Continual Learning Problem
IFL is framed as a sequential task problem where training sets arrive in an ordered stream (D[1] to D[Q]). The goal is for the model to perform well on all previous tasks while progressively learning to localize for newly arriving, unseen forgery domains without catastrophic forgetting.

Terminology

Summary

The rapid evolution of image manipulation techniques has raised pressing public security concerns, and existing Image Forgery Localization (IFL) methods often fail to adapt dynamically to newly emerging forgeries in real-world sequential data streams. This work introduces a novel continual learning framework designed to enable IFL models to progressively adapt to new forgery domains while retaining previously acquired forensic knowledge.

The gist: A forensic-aware continual adaptation framework is proposed, comprising a spatial mixture-of-forensic-experts (SMoFE) module for representation adaptation and a Fisher-weighted LoRA Gradient (FLAG) surgery strategy for knowledge preservation.

Problem Formulation and Benchmark

The paper reformulates Image Forgery Localization (IFL) as a practical continual learning problem, considering an ordered stream of sequential tasks, denoted as D[1] to D[Q], where training sets arrive sequentially. The authors establish a comprehensive benchmark under two realistic data evolution protocols: Protocol 1, which simulates sequential dataset arrival across four manipulation datasets (Classic, DEFACTO, FantasticReality, TampCOCO), and Protocol 2, which covers cross-content continual learning involving natural images, document images, and scientific images. Extensive evaluations of state-of-the-art IFL methods reveal substantial performance degradation under both settings, exposing two key challenges: (1) how to adaptively capture intrinsic forensic traces from incoming data across diverse unseen domains, and (2) how to preserve previously acquired forensic knowledge during sequential adaptation.

Forensic-Aware Representation Adaptation (FARA)

To address the challenge of adapting forensic representations, the framework introduces a two-part module:

  1. Spatial Mixture-of-Forensic-Experts (SMoFE): This mechanism performs fine-grained spatial routing over complementary forensic traces by adaptively combining local anomalies from different experts. The model utilizes four complementary trace extractors: E =lbrace Sobel, SRM, JPEG, Bayer… that capture edge discontinuities, high-frequency residuals, compression blockiness, and CFA/Bayer-pattern inconsistencies. A gate network predicts expert probabilities conditioned on both semantic visual features and forensic trace features to select the most relevant cues. The fused forensic trace embedding is computed as: z for = X sum K k=1 π˜ k ⊙ z k.

  2. Forensic Evidence-Guided Dense Prompting (FEGDP): This module addresses the forensic-to-prompt representation gap by transforming low-level forensic traces into structured localization evidence for SAM. It decomposes forgery evidence into two complementary cues: regional inconsistency, which identifies suspicious regions exhibiting abnormal forensic statistics (z inc), and boundary evidence, which captures transition patterns between manipulated and pristine regions (b hat). The final dense prompt is constructed as p = Hpr [z inc, bˆ, u], where u quantifies the normalized entropy of the expert probabilities, explicitly exposing ambiguity to the prompt composer.

Fisher-Weighted LoRA Gradient Surgery (FLAG)

To mitigate catastrophic forgetting and preserve knowledge during adaptation, the paper introduces a Fisher-weighted LoRA Gradient (FLAG) surgery strategy. This strategy accounts for the asymmetry in how different LoRA dimensions contribute to prior knowledge by estimating importance using diagonal Fisher information. After each task, its Fisher importance is estimated as Fi = EB∼Mi h(∇ϕL(B; Θ))2 i, and accumulated weights w<t = 1/t-1 sum Pt-1 i=1 Fi. This establishes an importance-aware geometry over the LoRA adaptation space. The surgery then selectively suppresses conflicting updates: g¯t = if g t, ⟨g t, g<t⟩w < 0, then gt -⟨gt, g<t⟩w/g<t2 w +ϵ else gt. This ensures that conflicts along old-task-critical LoRA dimensions are emphasized, thereby balancing stability and plasticity.

Overall Training Objective

The overall training objective L is a composite loss function: L = Lmask + λclsLcls + λbdLbd + λmoeLmoe, where Lmask and Lcls compute binary cross-entropy losses for mask prediction and image classification, respectively. An aggregated manipulation score yˆ is predicted by combining global visual and forensic features with evidence-map statistics: yˆ = C [z img, z for, Stat(z inc, bˆ, p)]. Boundary loss (Lbd) supervises the predicted boundary b̂ with the ground-truth boundary b. Finally, an expert balance loss (Lmoe) prevents the Top-1 router from collapsing to a single forensic expert.

Experimental Results and Contributions

Extensive experiments demonstrate that FOCAL achieves "state-of-the-art performance in both pixel-level forgery localization and image-level forgery detection across two continual learning protocols.

Improvements for AI systems

Based on the provided scientific paper, here are the specific improvements that can be implemented in AI systems (specifically Image Forgery Localization models) and what those improved systems will be capable of doing:


The proposed framework is called FOCAL (Forensic-Aware Continual Adaptation with LoRA surgery). The improvements focus on addressing the dual challenges of adapting to new forgery domains while preventing catastrophic forgetting.

Here are the specific enhancements:

  1. A novel forensic representation adaptation module, FARA, consisting of two components:

  2. Spatial Mixture-of-Forensic-Experts (SMoFE): This mechanism adaptively selects informative forensic traces from multiple heterogeneous experts (Sobel, SRM, JPEG, Bayer) at each spatial location using an element-wise MoE.

  3. Forensic Evidence-Guided Dense Prompting (FEGDP): This module transforms the low-level forensic traces into structured localization evidence for the Segment Anything Model (SAM) backbone by predicting:

  4. Inconsistency Evidence Maps: Highlighting regions with abnormal local feature statistics across multiple scales to identify suspicious areas.

  5. Boundary Evidence Maps: Modeling transition patterns between manipulated and pristine regions by jointly fusing forensic, visual embeddings, and inconsistency maps to distinguish genuine boundaries from benign edges.

  6. A Fisher-weighted LoRA Gradient Surgery (FLAG) strategy for knowledge preservation: This module implements an importance-aware geometry over the LoRA adaptation space by estimating the diagonal Fisher information of previous task gradients.

  7. Importance-Aware Gradient Surgery: It selectively suppresses gradient updates that conflict with directions critical to old tasks (identified via the Fisher-weighted inner product), thereby mitigating catastrophic forgetting while allowing plasticity for new domains.

The improved AI system (FOCAL) can perform the following specific capabilities:

  1. Adapt to Sequential Forgery Data Streams: The system can process a continuous, sequential stream of images arriving from different datasets (cross-dataset protocol) or different content types (cross-content protocol), such as moving from classic manipulations to scientific or document image forgeries.

  2. Robust Localization in Unseen Domains: Unlike standard IFL methods that degrade significantly when encountering novel forgery techniques or domain shifts, the FOCAL system can maintain high performance on previously learned manipulation patterns while effectively learning and localizing artifacts in entirely new, unseen forgery domains.

  3. Fine-Grained Spatial Forensic Analysis: The SMoFE mechanism allows the model to dynamically capture heterogeneous low-level traces (e.g., compression artifacts vs. noise residuals) at every pixel, enabling highly localized detection of subtle manipulation cues that might be missed by single-feature detectors or standard MoE structures.

  4. Structured Evidence Generation for State-of-the-Art Models: FEGDP converts raw, unstructured forensic data into a dense prompt format (inconsistency map, boundary map, and routing uncertainty) specifically tailored for the SAM backbone. This allows the system to leverage the powerful semantic understanding of SAM while providing it with explicit, forensic evidence rather than relying solely on generic visual features or simple prompts.

  5. Stable Continual Learning Performance: The FLAG surgery strategy ensures that as the model adapts to new tasks, it explicitly protects the parameter subspace responsible for localization knowledge from being overwritten by conflicting updates from new domains, leading to superior stability-plasticity trade-off compared to methods that suffer from catastrophic forgetting (e.g., SAM-B&LoRA or baseline IFL models).

  6. Generalization Across Image Content: By explicitly modeling cross-content evolution (Natural, Document, Scientific images), the system demonstrates strong robustness in handling large domain gaps between sequential tasks—a critical failure point for existing IFL methods.

Abstract

The rapid evolution of image manipulation techniques has raised growing public security concerns. Existing Image Forgery Localization (IFL) methods can accurately localize manipulated regions but are often unable to adapt to newly emerging forgeries. In real-world forensic scenarios, data typically arrive sequentially, yet continual model adaptation remains largely unexplored in IFL. To bridge this gap, we introduce the first continual learning framework for IFL and establish a comprehensive benchmark under two realistic data-evolution protocols: cross-dataset and cross-content continual learning. Evaluations of representative state-of-the-art IFL and continual learning methods reveal substantial performance degradation, highlighting two key challenges: (1) adaptively capturing intrinsic forensic traces from incoming data across unseen domains, and (2) preserving previously acquired forensic knowledge during sequential adaptation. To address these challenges, we propose a forensic-aware continual adaptation framework. First, a forensic trace mining module employs Spatial Mixture-of-Forensic-Experts (SMoFE) to dynamically route complementary forensic cues across spatial locations, together with Forensic Evidence-Guided Dense Prompting (FEGDP) to transform low-level forensic traces into structured localization evidence for SAM. Second, Fisher-weighted LoRA Gradient (FLAG) surgery identifies old-task-sensitive adaptation directions and suppresses conflicting updates, mitigating catastrophic forgetting while preserving plasticity for emerging forgery domains. Extensive experiments demonstrate state-of-the-art performance in both pixel-level forgery localization and image-level forgery detection across diverse continual learning scenarios.

Related papers