Guides That Cause Actions: An Offline Study of Guide-Action Mutual Reinforcement in Multimodal Web Agents
cs.CL
Submitted: 2026-09-23
Updated: 2026-09-23
Terminology
Sources
- GUI-R1 : A Generalist R1-Style Vision-Language Action Model For GUI Agents
- MAG: A Web-Agent Benchmark and Harness for Multimodal Action and Guide Generation
- Understanding R1-Zero-Like Training: A Critical Perspective
- WebLINX: Real-World Website Navigation with Multi-Turn Dialogue
- UI-TARS: Pioneering Automated GUI Interaction with Native Agents
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- AgentTrek: Agent Trajectory Synthesis via Guiding Replay with Web Tutorials
- Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering