SynWeaver: Website-Prior Task and Trajectory Co-Synthesis for Web Agents
Ruitao Wang, Yuwen Hao, Menglin Yang
Hong Kong University of Science and Technology (Guangzhou)
cs.SE, cs.AI
Submitted: 2026-08-12
Updated: 2026-08-14
Comments: 31 pages, 9 figures
Code: https://github.com/Eilok/SynWeaver
Project page: https://storage.googleapis.com/deepmind-media/ModelCards/Gemini-3-Flash-Model-Card.pdf
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 75/100
The gist: SynWeaver is a website-prior task and trajectory co-synthesis framework designed to address two key limitations in existing exploration-based data synthesis methods for web agents: (1) they often
Terminology
Summary
SynWeaver is a website-prior task and trajectory co-synthesis framework designed to address two key limitations in existing exploration-based data synthesis methods for web agents: (1) they often fail to cover the full functionality of a website, and (2) without sufficient website prior knowledge, they tend to propose hallucinated tasks, which limits the diversity and efficiency of downstream trajectory synthesis.
The framework proceeds in three stages. First, it constructs a website map using a depth-first search (DFS)-based crawler. This map captures functionally distinct page states and executable interactions, reducing redundant exploration compared to random-walk approaches. The paper states: "We introduce a DFS-based crawler that constructs a website map G = (S, T) for each target website, where each node s ∈ S denotes a functionally distinct state together with its page screenshot and accessibility tree, and each edge τ ∈ T denotes an executable interaction and records the corresponding action." The crawler uses a unified state processor with progressive state comparison and duplicate-trigger detection to avoid over-exploring superficial page variations.
Second, SynWeaver performs website-prior learning. It derives five types of supervision from the website map: page description, page question answering, element description, forward transition description, and inverse transition description. These are used to fine-tune a UI-aware model via LoRA. The paper notes: We therefore learn website-specific UI priors from the website map before task synthesis.
This enables more grounded task proposals.
Third, SynWeaver performs collaborative task-trajectory synthesis. Unlike prior methods that treat task refinement and trajectory refinement as separate problems, SynWeaver jointly updates the task and execution trajectory when they become inconsistent. The paper explains: We trigger collaborative refinement whenever the current task lacks essential details, becomes incompatible with the observed website state, or execution stalls after multiple unsuccessful attempts.
Two refinement modes are used: task-only mode (updating the task while keeping the trajectory fixed) and joint mode (editing both task and trajectory, including deleting or reordering steps and updating reasoning). After execution, a post-verification stage applies heuristic checks and, if needed, reconstruction to produce executable, semantically aligned supervision.
Experiments were conducted on WebArena and WebVoyager. On WebArena, using only 822 validated task-trajectory pairs, SynWeaver achieved the best overall success rate on both Qwen3-VL-8B-Instruct (19.91%) and InternVL3-8B (14.16%), outperforming the strongest baseline SynthAgent by 3.10 and 1.33 points respectively. On WebVoyager, SynWeaver reached 27.06% success rate, outperforming SynthAgent by 4.64 points. The paper states: "These gains are noteworthy because, under our matched training setup, all methods are fine-tuned from the same UI-aware initialization Mui. This indicates that the advantage mainly comes from the quality of the synthesized task-trajectory pairs rather than from differences in model initialization."
Ablation studies showed that removing the website map (−Map) reduced overall success rate by 4.42 points, replacing collaborative refinement with decoupled refinement (−CR) caused a 4.87-point drop, replacing the website-prior proposer with a general-purpose model (−WP) lowered performance by 2.21 points, and removing post-verification (−PV) produced the largest overall drop of 5.31 points.
The paper also demonstrated that website maps improve task diversity, with SynWeaver achieving the best NovelSum scores across all three settings (0.3805, 0.3731, and 0.3854). Additionally, collaborative refinement improved task success rate from 71.89% to 74.79% and trajectory retention from 90.59% to 99.52%, while reducing synthesis cost from 0.091 to 0.085.
The contributions are summarized as: (1) a structured website exploration mechanism that reduces redundant interactions while preserving coverage of functionally distinct states; (2) SynWeaver, a website-prior-centered framework for collaborative task-trajectory synthesis; and (3) extensive experiments showing more effective and data-efficient supervision for web agents, yielding strong in-domain and out-of-domain gains over prior synthesis methods.
Improvements for AI systems
Improvements to AI Systems:
- Website-Aware Task Proposer with Grounded Priors
-
The AI system can generate web-agent tasks that are guaranteed to be executable on a given site by first building a structured map of states and interactions (via DFS crawling).
-
It avoids hallucinated or infeasible tasks by conditioning task proposals on learned website-specific UI priors (page descriptions, element/transition QA, forward/inverse transitions).
-
Capability: Automatically propose diverse, realistic user goals (e.g.,
book a flight with two stops
orchange account email
) that match the actual site’s functionality, not generic or impossible requests.
- Collaborative Task–Trajectory Refinement
-
The AI system can jointly revise both the task description and the execution trajectory when they become inconsistent (e.g., task lacks detail, state mismatch, or repeated failures).
-
It supports two modes: task-only updates (keeping trajectory fixed) and joint edits (deleting/reordering steps, updating reasoning).
-
Capability: Self-correct during data synthesis, producing higher-quality supervision pairs where the task and the action sequence are perfectly aligned—reducing noise that would otherwise confuse downstream agent training.
- Efficient Exploration with Progressive State Deduplication
-
The system uses a unified state processor that compares page states progressively and detects duplicate triggers, avoiding redundant crawling of superficial page variations (e.g., minor DOM changes).
-
Capability: Build a compact but functionally complete website map with fewer interactions, lowering data collection cost and time while preserving coverage of all distinct user-reachable states.
- Post-Verification and Reconstruction for Reliable Supervision
-
After synthesis, the system applies heuristic checks (e.g., action validity, state reachability) and, if needed, reconstructs the trajectory to ensure every step is executable and semantically aligned with the task.
-
Capability: Output only validated task–trajectory pairs, preventing the propagation of broken or misleading examples into the training set—leading to more robust web agents.
- Data-Efficient Fine-Tuning of UI-Aware Models
-
The system fine-tunes a UI-aware model (via LoRA) on website-specific priors before task synthesis, then uses the synthesized pairs for downstream agent training.
-
Capability: Achieve strong web-agent performance with only 800 validated pairs per site (as shown on WebArena), reducing the need for large human-annotated datasets while improving in-domain and out-of-domain generalization.
- Diverse Task Generation with Novelty Scoring
-
The system’s map-based exploration naturally leads to higher task diversity (measured by NovelSum), enabling the agent to handle a wider range of user intents rather than overfitting to common patterns.
-
Capability: Train web agents that generalize better to unseen tasks and websites, as they have seen a broader distribution of goals and interaction sequences.
- Cost-Aware Synthesis Pipeline
-
By reducing redundant exploration and using collaborative refinement (which increases trajectory retention from 90% to 99.5%), the system lowers synthesis cost (e.g., from 0.091 to 0.085 per pair).
-
Capability: Scale up data generation for web agents at lower expense, making it feasible to cover many websites or frequent site updates without prohibitive API costs.
Sources
- Qwen3-VL Technical Report
- Go-Browse: Training Web Agents with Structured Exploration
- WebSynthesis: World-Model-Guided MCTS for Efficient WebUI-Trajectory Synthesis
- WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models
- UI-Voyager: A Self-Evolving GUI Agent Learning via Failed Experience
- Structured Distillation of Web Agent Capabilities Enables Generalization
- NNetNav: Unsupervised Learning of Browser Agents Through Environment Interaction in the Wild
- A Survey on (M)LLM-Based GUI Agents
- UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning
- SynthAgent: Adapting Web Agents with Synthetic Supervision
- UI-TARS: Pioneering Automated GUI Interaction with Native Agents
- AutoWebWorld: Synthesizing Infinite Verifiable Web Environments via Finite State Machines
- ScribeAgent: Towards Specialized Web Agents Using Production-Scale Workflow Data
- GUI Knowledge Bench: Revealing the Knowledge Gap of VLMs in GUI Tasks
- Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V
- WebNavigator: Global Web Navigation via Interaction Graph Retrieval
- InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Related papers
- Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
- GitSkills: A Dataset of Agent Skills on GitHub
- SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
- PackMonitor: Enabling Zero Package Hallucinations Through Decoding-Time Monitoring
- IntentCoding: Amplifying User Intent in Code Generation
- Incentives and Outcomes in Bug Bounties