Paint What You See: Benchmarking Dexterous Visual Tool Use in Multimodal Agents
cs.AI
Submitted: 2026-08-26
Updated: 2026-08-26
Code: https://github.com/OOOHS/EASEL
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Qwen3-VL Technical Report
- Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic
- Gemma 3 Technical Report
- GLM-5V-Turbo: Toward a Native Foundation Model for Multimodal Agents
- DeepEyesV2: Toward Agentic Multimodal Model
- Kimi K2.5: Visual Agentic Intelligence
- Imagine while Reasoning in Space: Multimodal Visualization-of-Thought
- Step1X-Edit: A Practical Framework for General Image Editing
- MiMo-VL Technical Report
- Towards Pixel-Level VLM Perception via Simple Points Prediction
- AgentVista: Evaluating Multimodal Agents in Ultra-Challenging Realistic Visual Scenarios
- DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding
- Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V
- Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement
- Thyme: Think Beyond Images
- DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning
- InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection