SPARED: Reasoning-Based AI-Generated Image Detection via Adversarially Edited Data

arXiv:2608.12876 · cs.CV, cs.AI · Submitted 2026-08-13 · Read on arXiv

Yicheng Bao, Xiahui Guo, Xuhong Wang, Xin Tan

East China Normal University · Shanghai Artificial Intelligence Laboratory

cs.CV, cs.AI

Submitted: 2026-08-13

Updated: 2026-08-14

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 95/100

The gist: SPARED introduces an adversarial reinforcement learning framework for AI-generated image detection that pits two heterogeneous models against each other: a diffusion image editor that learns to edit

Terminology

Summary

SPARED introduces an adversarial reinforcement learning framework for AI-generated image detection that pits two heterogeneous models against each other: a diffusion image editor that learns to edit real photographs into fake counterparts that fool the current detector, and a reasoning multimodal large language model (MLLM) that learns to expose them with a verdict grounded in free-form reasoning. The framework addresses three failure modes inherited by existing detectors: provenance shortcuts from real and fake images collected from different sources, templated rationales from supervised explanation corpora, and a static forgery corpus that leaves the decision boundary stationary while generators keep moving.

The defender is built on a Qwen3.5-9B multimodal backbone, first LoRA-tuned on real/fake reasoning pairs from the DeepfakeJudge training corpus to teach the tag format and elementary artifact vocabulary, then continued with full-parameter GRPO. The defender generates an open-ended reasoning span followed by an explicit verdict, and only the final verdict is rewarded—explanations cannot score by reproducing templates and improve only insofar as they lead to correct verdicts. The attacker is a LoRA adapter on Qwen-Image-Edit-2511 trained online with DiffusionNFT, and it edits real source photographs into their fake counterparts rather than synthesizing fakes from scratch, so every fake shares near-identical content, composition, and resolution with its real source, leaving only the residual editing trace for the defender to learn from.

The attacker's reward is gated: it is credited only when an instruction-following gate (PaCo) confirms the edit was faithfully executed and the resulting pair defeats the current defender, so fooling without genuinely editing earns nothing. The two models are architecturally heterogeneous—a diffusion editor in continuous pixel space against an autoregressive MLLM over discrete tokens—and share no parameters. The training schedule alternates the two models over five rounds: the first defender GRPO round trains on a pool edited by the base image-editing model, each attacker round trains against a frozen snapshot of the preceding defender, and the next defender round continues GRPO from the previous defender checkpoint on a pool regenerated by the newly trained attacker.

Every fake in every round comes from editing a real image that is also present as that pair's real label, and no source image may leak into the images used to build evaluation benchmarks. A fixed source set is randomly sampled from a mixed single-turn editing corpus (ImgEdit, pico-banana-400k, MagicBrush), deduplicated, and screened by perceptual-hash matching against every benchmark used in subsequent evaluation.

The detector trained within this loop improves monotonically across rounds on three external benchmarks. On DeepfakeJudge-Detect, overall accuracy rises from 69.6 through SFT to 72.1 at Iter1, 76.0 at Iter2, and 79.5 at Iter3, with real recall climbing from 57.1 to 70.2 and fake recall from 82.0 to 88.8. At 9B parameters, the final model surpasses every non-reasoning MLLM evaluated including Qwen3-VL-235B (74.5), and every reasoning model up to 30B (77.5), trailing only Qwen3-VL-235B-Thinking (82.7), a model 26 times its size. The gain concentrates in the locally edited subset of the fake pool, where accuracy rises from 44.8 to 82.5, matching a training pool built from such edits.

On AnomReason-Deepfake, SPARED reaches 92.18 accuracy zero-shot, above every closed model compared against (GPT-4o: 87.76), 3.6 points above the strongest baseline UniGenDet (88.59), and 9.6 points above AnomReasonor (82.61), although the latter is fine-tuned on this benchmark's training data. Explanation quality moves the same way: CSemAP-Full reaches 0.5207, a 23% relative improvement over the strongest baseline score (0.4234). The SFT stage trades accuracy for the reasoning format (84.05 → 77.83 Acc, but 0.3063 → 0.4792 CSemAP), and the subsequent GRPO rounds, whose reward never scores the explanation, lift both metrics simultaneously and monotonically (85.72/0.4927 → 87.58/0.5103 → 92.18/0.5207).

On Holmes-Set, which contains fully synthetic images from ten unseen generator families, mean accuracy rises monotonically from 53.0 (base) through 65.7 (SFT) and 84.5 (Iter1) to 92.8 (Iter3) with zero in-domain data, within 2.8 points of AIGI-Holmes (95.6). Adding in-domain detection data to Iter3 reaches 97.2 mean accuracy and 99.9 mean AP, the best AP in the table, closing most of the distance to the strongest specialists GenShield (98.8) and UniGenDet (99.2). The same mix given to the SFT model without adversarial rounds reaches 89.5, behind even the zero-shot Iter3, showing that in-domain data compounds the adversarial rounds rather than substituting for them.

Ablation studies confirm that the adversarial rounds are necessary: continued GRPO on the frozen round-one pool captures only 2.2 of the 7.4 points that the adversarial branch adds on DeepfakeJudge-Detect, and regenerating the pool each round from fresh sources with the frozen base editor peaks at 73.7 and oscillates below that thereafter. Removing the instruction-fidelity gate causes the attacker to pass the fidelity check 9.6 points less often, shrink its median edit magnitude by 27%, and fool the frozen defender slightly more while editing less; training a defender round on this pool falls below its starting point (66.7 vs. Iter1's 72.1 on DeepfakeJudge-Detect; 81.12 vs. 85.72 on AnomReason-Deepfake), with real recall collapsing to 38.2 while fake recall inflates to 95.1. Unpairing the training pool—replacing the reals with photographs from disjoint sources—causes most of the round's benefit to disappear: the unpaired round gains 1.1 points on DeepfakeJudge-Detect where the paired round gains 3.9, and on AnomReason-Deepfake it lands below even its starting point.

The main limitation identified is the hard 0/1 fooling reward: without per-family difficulty control, an isolated per-generator regression persists even as the means rise. Specifically, Janus accuracy drops from 86.5 (Iter2) to 73.1 (Iter3) even as every other family's accuracy improves and the mean rises from 90.2 to 92.8; with in-domain data added, Janus recovers to 84.3. The paper concludes that a graded difficulty signal is the natural next step.

Improvements for AI systems

Improvements to AI systems based on this paper:

  1. Adversarial co-training for robust deepfake detection: Implement a dual-model loop where a generative editor (diffusion-based) and a reasoning detector (MLLM) are alternately trained. The editor learns to produce fakes that fool the current detector, while the detector learns to expose them via free-form reasoning. This yields monotonic accuracy gains (69.6→79.5 on DeepfakeJudge-Detect) and closes the gap to models 26× larger.

  2. Gated reward for faithful adversarial edits: Use an instruction-fidelity gate (e.g., PaCo) to reward the attacker only when edits are both faithful to instructions and fool the defender. This prevents degenerate shortcuts—without the gate, edit magnitude shrinks 27%, real recall collapses to 38.2, and defender performance drops below baseline.

  3. Paired real-fake training with content-matched sources: Ensure every fake is an edit of its paired real image, sharing identical content/composition/resolution. This removes provenance shortcuts and forces the detector to learn residual editing traces. Unpairing the pool reduces gains by 70% (1.1 vs 3.9 points on DeepfakeJudge-Detect).

  4. Reasoning-based verdicts with verdict-only reward: Train the detector to output open-ended reasoning followed by an explicit verdict, but reward only the verdict’s correctness. This improves both accuracy and explanation quality (CSemAP 0.3063→0.5207) without templated rationales, and avoids overfitting to supervised explanation corpora.

  5. Dynamic regeneration of adversarial training pools: Regenerate the fake pool each round using the latest attacker, rather than freezing a static corpus. Static pools capture only 2.2 of 7.4 possible points; dynamic regeneration with fresh sources peaks at 73.7 and oscillates, while full adversarial rounds reach 79.5.

  6. Zero-shot generalization to unseen generator families: Train exclusively on edited real photos, then evaluate on fully synthetic images from unseen generators. This yields 92.8% accuracy on Holmes-Set (vs 53.0 base) with zero in-domain data, showing that editing-based training transfers to synthetic detection.

  7. Graded difficulty signal for per-family robustness: Replace the hard 0/1 fooling reward with a graded signal that accounts for per-generator difficulty. This addresses the observed regression (Janus drops 86.5→73.1) while mean accuracy rises, preventing isolated failures in adversarial loops.

What the improved AI system can do:

  • Detect AI-generated images with higher accuracy and more interpretable, free-form explanations than current state-of-the-art, even at 9B parameters (surpassing 235B non-reasoning models).

  • Adapt continuously to new generative models without manual corpus updates, maintaining performance on unseen generators.

  • Provide trustworthy verdicts with reasoning that humans can audit, while avoiding both provenance shortcuts and templated rationales.

  • Achieve strong zero-shot performance on fully synthetic content, and near-specialist performance when given small in-domain data (97.2% accuracy, 99.9% AP).

  • Maintain balanced precision/recall (real recall 70.2, fake recall 88.8) rather than collapsing to trivial solutions, thanks to gated adversarial training.

Sources

Related papers