VLM-in-Sandbox: Visual Workspaces for Agentic Visual Reasoning
cs.AI
Submitted: 2026-09-21
Updated: 2026-09-21
License: http://creativecommons.org/licenses/by/4.0/
The gist: Sandboxed computer environments support multi-step reasoning with tools, executable programs, and persistent files, yet their extension from language models to vision-language models (VLMs)
Terminology
Abstract
Sandboxed computer environments support multi-step reasoning with tools, executable programs, and persistent files, yet their extension from language models to vision-language models (VLMs) introduces a distinct state-management problem. Visual reasoning produces intermediate image-valued evidence---crops, masks, overlays, zoomed regions, and analytic renderings---that must remain addressable without accumulating unboundedly in multimodal context. We introduce VLM-in-Sandbox, a training-free framework for agentic multimodal reasoning in controlled computer environments. Its Visual Workspace registers generated artifacts in an image ledger, maintains a bounded active visual context, and lets the model explicitly promote selected evidence for subsequent inspection. This separates visual evidence generation, performed by sandbox tools, from visual evidence management. Across seven benchmarks and four base VLMs, VLM-in-Sandbox achieves the highest sample-weighted average accuracy among Vanilla VLM, Append-only Sandbox, and the proposed method. A compiler-matched 2 times2 study on 1,260 examples further separates model-directed visibility from bounded retention: VLM-in-Sandbox reaches 66.27% accuracy with 18.6% fewer total tokens than the automatic, retain-all control. Over all 6,350 submitted GPT-4.1-mini examples, it produces 302 rescues and 142 regressions relative to Original Append-only. A local vLLM study with prefix caching confirms that the smaller request workload also reduces uncached tokens, time to first token, and end-to-end latency. These results identify explicit visual evidence state as a central abstraction for sandboxed VLM agents.
Sources
- Are We on the Right Way for Evaluating Large Vision-Language Models?
- Computer Environments Elicit General Agentic Intelligence in LLMs
- ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding
- Visual Programming: Compositional visual reasoning without training
- Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models
- Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models
- VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks
- TIR-Bench: A Comprehensive Benchmark for Agentic Thinking-with-Images Reasoning
- Beyond Accuracy: Evaluating Grounded Visual Evidence in Thinking with Images
- Abstract 3D Perception for Spatial Intelligence in Vision-Language Models
- Visual-RFT: Visual Reinforcement Fine-Tuning
- MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
- Foveated Reasoning: Stateful, Action-based Visual Focusing for Vision-Language Models
- We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?
- CodeDance: A Dynamic Tool-integrated MLLM for Executable Visual Reasoning
- OpenThinkIMG: Learning to Think with Images via Visual Tool Reinforcement Learning
- ViperGPT: Visual Inference via Python Execution for Reasoning
- Reason-RFT: Reinforcement Fine-Tuning for Visual Reasoning of Vision Language Models
- LlamaV-o1: Rethinking Step-by-step Visual Reasoning in LLMs
- Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection