MAD-Guard: Controlled Study of Autoregressive Generation versus Direct Decision Interfaces for Closed Multimodal Forensic Tasks
cs.CV, cs.CR
Submitted: 2026-09-27
Updated: 2026-09-27
Code: https://github.com/moyuan10086/MAD-Guard
Terminology
Sources
- Deep Residual Injection for Full-Spectrum Forensic Signal Perception in Multimodal Large Language Models
- PRPO: Paragraph-level Policy Optimization for Vision-Language Deepfake Detection
- VIGIL: Part-Grounded Structured Reasoning for Generalizable Deepfake Detection
- FakeVLM-R1: Internalizing Physical Laws via CoT for Synthetic Image Detection
- Perception, Verdict, and Evolution: Hindsight-Driven Self-Refining Forensics Agent for AI-Generated Image Detection
- Evidence-Grounded Forensic Reasoning for Detecting and Grounding Multi-Modal Media Manipulation
- Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
- Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
- GLiNER Guard: Unified Encoder Family for Production LLM Safety and Privacy
- SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning
- Training Deep Nets with Sublinear Memory Cost
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models