ScreenHaystack: Finding Blind Zones in GUI Grounding
cs.CV
Submitted: 2026-09-25
Updated: 2026-09-25
Code: https://github.com/mlfoundations/gelato
Terminology
Sources
- Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
- Qwen3-VL Technical Report
- UI-Venus Technical Report: Building High-performance UI Agents with RFT
- UI-Vision: A Desktop-centric GUI Benchmark for Visual Perception and Interaction
- Improved GUI Grounding via Iterative Narrowing
- UI-TARS: Pioneering Automated GUI Interaction with Native Agents
- Learning from Failure: Inference-Time Self-Improvement for Computer-Use Agents
- MMBench-GUI: Hierarchical Multi-Platform Evaluation Framework for GUI Agents
- GUI-Perturbed: Domain Randomization Reveals Systematic Brittleness in GUI Grounding Models
- Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models