Dual-Modality Multi-Stage Adversarial Safety Training: Robustifying Multimodal Web Agents Against Cross-Modal Attacks
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Dual-Modality Multi-Stage Adversarial Safety Training".
Jane: Multimodal web agents that process both screenshots and accessibility trees are increasingly deployed to interact with web interfaces, yet their dual-stream architecture opens an underexplored attack surface:
Tom: First, who's behind it and why it matters.
Paper summary: Tom: So, to recap, this paper is all about how multimodal web agents that process both screenshots and accessibility trees are vulnerable because they have two observation channels that an attacker can manipulate simultaneously with a consistent deceptive narrative.
Jane: The paper claims that attacks including visual components perform significantly better than text-only injections when targeting these multimodal agents, which exposes weaknesses in safety training methods focused only on text.
Lu: Essentially, the research formalizes the agent–attacker interaction as a two-player zero-sum Markov game to rigorously describe this relationship and then proposes DMAST as a three-stage pipeline to address it.
Meng: The core idea seems to be using adversarial co-evolution, where both the agent and attacker are trained together from the same underlying VLM with shared weights.
Lalam: It’s about building a method that forces the models to learn robust behaviors by having them actively compete against each other in a structured training setup.
Tom: And this process involves imitation learning, followed by an oracle-guided fine-tuning step where the agent learns task-focused reasoning under adversarial noise without acknowledging the attack during that phase.
Jane: The final stage uses Adversarial Reinforcement Learning through GRPO self-play, where the agent and attacker iterate against each other to refine their policies.
Lu: It’s a comprehensive approach because it moves beyond simple defense by co-training the players in this dynamic game to create defenses that can handle coordinated visual and textual deception.
Meng: I see how that multi-stage approach is necessary; you need a stable base before you can effectively start pitting the models against each other in the adversarial setting.
Lalam: I think this framework’s main contribution is showing how to systematically harden multimodal agents against cross-modal attacks through this structured, iterative training pipeline.
Conclusion: Tom: So, looking at the title, "Dual-Modality Multi-Stage Adversarial Safety Training: Robustifying Multimodal Web Agents Against Cross-Modal Attacks," it clearly tells us this work is focused on strengthening how these agents handle complex threats across multiple observation types.
Jane: The authors are Haoyu Liu, Dingcheng Li, Lukas Rutishauser, and Zeyu Zheng from UC Berkeley and Google Deepmind; their research focuses on formalizing the interaction as a game to create a structured training methodology.
Lu: What this means in simpler terms is that we’re moving toward agents that aren't just reacting; they are actively trained to anticipate and counter coordinated attacks across different data streams.
Meng: The implication for me is that for practical AI deployment, this suggests we need to prioritize safety training methods that explicitly account for cross-modal risks, especially when dealing with visual inputs.
Lalam: I think the broader impact is establishing a new standard of robustness where agents are trained against sophisticated threats in a structured way rather than just hoping they won't fail in novel scenarios.
Tom: It really moves the concept from reactive safety measures to proactive, co-evolutionary training, suggesting that this approach offers better performance when agents encounter tricky, unseen web environments.
Jane: Ultimately, the paper suggests that by employing this DMAST framework with prompt-based defense, we can achieve a lower Attack Success Rate across various benchmarks while keeping task success rates high.
Lu: This points toward a future where multimodal AI systems are inherently designed to be resilient against coordinated visual and textual manipulations during their operation.
Meng: It means we need to build in this kind of adversarial training early on, making it a fundamental part of the development cycle for any agent interacting with web interfaces.
UC Berkeley · Google DeepMind
cs.LG, cs.AI, cs.CL
Submitted: 2026-03-04
Updated: 2026-10-05
Importance score: 85/100
The gist: Multimodal web agents that process both screenshots and accessibility trees are increasingly deployed to interact with web interfaces, yet their dual-stream architecture opens an underexplored attack
Key concepts
- DualModality
- This refers to multimodal web agents that process two types of input: visual data like screenshots and structural data like accessibility trees. The vulnerability lies in the fact that an attacker can inject deceptive content into the webpage's structure, corrupting both observation channels at once, making it harder for the agent to distinguish truth from manipulation.
- DMAST
- This is a three-stage framework designed to train multimodal agents against attacks. It involves imitation learning from experts, supervised fine-tuning using an 'oracle' that prevents acknowledgment of attacks, and adversarial reinforcement learning where the agent and attacker play against each other to improve robustness.
- Zero-Acknowledgment Strategy
- This is a technique used in the training stage where an oracle generates reasoning (Chain-of-Thought) that is strictly focused on the task elements. Crucially, this reasoning does not acknowledge or react to the presence of an adversarial attack, ensuring the agent learns task focus regardless of visual or structural deception.
- Group Relative Policy Optimization (GRPO)
- This is a reinforcement learning method used in Stage 3 where both the agent and attacker are trained simultaneously against each other. It functions as a self-play system where policies are updated based on rewards for task success and penalties for leakage, leading to genuine co-evolution of both roles.
Terminology
Summary
Multimodal web agents that process both screenshots and accessibility trees are increasingly deployed to interact with web interfaces, yet their dual-stream architecture opens an underexplored attack surface: an adversary who injects content into the webpage DOM simultaneously corrupts both observation channels with a consistent deceptive narrative.
The gist
Attacks including a visual component far outperform text-only injections when targeting multimodal web agents.
Problem Formulation and Vulnerability Analysis
The research formalizes the agent–attacker interaction as a two-player zero-sum Markov game where the agent receives task descriptions and synthetic user data, aiming to complete tasks while avoiding sensitive data leakage. The attacker generates structured HTML modifications executed in the browser runtime, simultaneously corrupting both modalities. A systematic vulnerability assessment on MiniWob++ reveals that attacks including a visual component (either image-only or coordinated dual-modality) are significantly more effective at inducing safety jailbreaks than text-only injections. Specifically, image-only attacks achieved an attacker success rate of 34.4%, compared to 24.1% for text-only injections, exposing critical gaps in text-centric VLM safety training against visual deceptions like typographic overlays or fake system dialogs.
DualModality Multi-Stage Adversarial Safety Training (DMAST)
DMAST is a framework designed to harden multimodal web agents through adversarial co-evolution, formalizing the interaction as a two-player zero-sum Markov game and utilizing a three-stage pipeline. The training procedure involves:
-
Imitation Learning from a strong teacher model to distill expert trajectories into a smaller student model.
-
Oracle-Guided Supervised Fine-Tuning that uses a
novel zero-acknowledgment strategy
to instill task-focused reasoning under adversarial noise, where an oracle generates Chain-of-Thought (CoT) reasoning strictly grounded in task elements without acknowledging the attack. -
Adversarial Reinforcement Learning via Group Relative Policy Optimization (GRPO) self-play, co-evolving both the agent and attacker through iterative training against each other.
Training Pipeline Details
The DMAST framework is built upon a shared-weight design where both agent and attacker are instantiated from the same VLM with role-specific system prompts to enable efficient co-evolution.
(Stage 1: Imitation Learning)
This stage collects demonstration data, including adversarial episodes (agent succeeds despite attacks) and clean episodes, which are then used to fine-tune the student model via a KL-regularized SFT objective.
(Stage 2: Oracle-Guided SFT)
This stage constructs a synthetic denoising
dataset by generating golden trajectories (clean task completions) and corresponding attacked versions where an oracle generates task-focused CoT reasoning that adheres to the zero-acknowledgment principle,
ensuring the target action remains invariant while the observation changes.
(Stage 3: Adversarial RL)
This stage employs Group Relative Policy Optimization (GRPO) self-play. The agent and attacker interact in a Markov game, with rewards incentivizing task completion while penalizing data leakage. The policy is updated via a clipped surrogate objective, and an asymmetric population strategy is used where the agent interacts with historical attacker checkpoints while the attacker trains against only the latest agent checkpoint.
Empirical Results and Co-evolutionary Dynamics
Experiments on held-out MiniWob++ tasks and out-of-distribution VisualWebArena tasks demonstrate DMAST's effectiveness. On curated VisualWebArena tasks, the full pipeline reduced attacker success from 41.2% (base model) to 21.4% while improving agent task completion from 6.2% to 10.2%. Cross-evaluation heatmaps reveal a classic adversarial arms race,
confirming genuine capability co-evolution: the Iteration 10 agent improved its success rate against a base attacker, and the Iteration 5 attacker tripled the base attacker’s success rate, demonstrating that both roles genuinely co-evolve through self-play. Furthermore, analysis of attack diversity metrics shows that lexical diversity (Distinct-n) and strategy entropy consistently increase across RL iterations, confirming that co-evolution naturally incentivizes attack diversification into more sophisticated, context-aware strategies.
Conclusion
DMAST establishes a superior balance between adversarial robustness and functional utility compared to established baselines like SPAG or Online SFT. The framework proves that the full pipeline consistently dominates at every RL iteration, confirming that Oracle-Guided SFT instills robust behavioral priors that RL optimizes rather than replaces, leading to agents that are both safer and more capable on complex, unseen environments. The combination of DMAST with prompt-based defense yields the lowest Attack Success Rate across both benchmarks while maintaining high Task Success Rate.
Improvements for AI systems
Here are the specific improvements that can be made to AI systems based on the Dual-Modality Multi-Stage Adversarial Safety Training (DMAST) framework, and what these improved systems can achieve:
-
Maturity of Safety Defense Against Cross-Modal Attacks:
-
Transition from Text-Centric to Visually Aware Safety Training:
-
Co-evolutionary Robustness Against Adaptive Adversaries:
-
Enhanced Task Completion Under Deceptive Conditions (Improved Utility):
5.1 Specific Improvements and Capabilities
The DMAST framework allows for the creation of multimodal web agents that possess significantly enhanced security and functional reliability compared to current state-of-the-art models (like base VLMs or standard prompt defenses).
Improvement Focus Mechanism Provided by DMAST Component Specific Capability Gained in AI System
:---:---:---
-
Cross-Modal Defense Robustness (Visual + Text) Dual-stream observation processing formalization and HTML injection mechanism. Stage 2 (Oracle-Guided SFT) enforces task focus by training the agent to ignore visually deceptive overlays, even when they appear in both the screenshot and the accessibility tree. The system can reliably complete complex tasks on web interfaces while being completely immune to sophisticated visual phishing, typographic overlays, or fake dialogs that are designed to mislead a purely text-based model.
-
Resilience Against Coordinated Attacks (Arms Race) Stage 3 (Adversarial RL via GRPO self-play). The agent and attacker share weights and co-evolve against each other over multiple iterations, forcing the agent to develop strategies that are robust not just to static attacks, but to dynamically evolving, context-aware injections. The system develops emergent defenses capable of recognizing multi-step coordinated attacks (e.g., instruction override followed by contextual phishing) and can resist novel injection techniques that have never been seen before in the training set.
-
Task-Focused Reasoning Under Noise Stage 2 (Oracle-Guided SFT) utilizing a privileged oracle model to generate Chain-of-Thought (CoT) that strictly adheres to the zero-acknowledgment principle. This teaches the agent to maintain goal adherence despite corrupted observations. The system demonstrates superior functional utility; it can execute precise, multi-step workflows accurately even when its visual input is heavily manipulated, ensuring high Task Success Rate (TSR) without succumbing to
refusal collapse.
-
Superior Safety-Utility Trade-off The entire three-stage pipeline (Imitation → Oracle SFT → RL) provides cumulative gains. The final RL stage optimizes for both minimizing Attack Success Rate (ASR↓) and maximizing Task Success Rate (TSR↑). The system achieves a
sweet spot
where it is highly robust against data leakage while maintaining high efficiency, outperforming existing methods like SPAG or pure prompt defenses on complex, unseen tasks (VisualWebArena). -
Adaptive Attack Generation Awareness Post-RL analysis shows the attacker evolves from template-like attacks to task-context-aware and multi-step coordinated strategies. The agent learns to predict and counter these evolving patterns by focusing on the underlying goal rather than surface anomalies. The system is capable of handling
zero-shot
or unseen adversarial scenarios because it has been trained via self-play against a diverse population of increasingly creative attack strategies, leading to better generalization to out-of-distribution (OOD) environments.
Sources
- Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models
- Towards Transferable Attacks Against Vision-LLMs in Autonomous Driving with Typography
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks
- Reinforcement Learning on Web Interfaces Using Workflow-Guided Exploration
- E$^2$AT: Multimodal Jailbreak Defense via Dynamic Joint Optimization for Multimodal Large Language Models
- Vision-LLMs Can Fool Themselves with Self-Generated Typographic Attacks
- AI Safety in Practice: Enhancing Adversarial Robustness in Multimodal Image Captioning
- Robust CLIP: Unsupervised Adversarial Fine-Tuning of Vision Embeddings for Robust Large Vision-Language Models
- Jailbreak in pieces: Compositional Adversarial Attacks on Multi-Modal Language Models
- Lessons from Defending Gemini Against Indirect Prompt Injections
- Gemma 3 Technical Report
- Typographic Attacks in a Multi-Image Setting
- Adversarial Reinforcement Learning for Large Language Model Agent Safety
- Self-Play Preference Optimization for Language Model Alignment
- AdvAgent: Controllable Blackbox Red-teaming on Web Agents
- Multimodal Prompt Injection Attacks: Risks and Defenses for Modern LLMs
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks