VTOS: Learning to Orchestrate Vision Tools by Co-Searching Solutions and Observers
cs.CV, cs.CL
Submitted: 2026-06-17
Updated: 2026-09-02
Comments: 19 pages, 6 figures, 9 tables. Accepted to EMNLP 2026 (Main Conference). Code: https://github.com/jinchaogjc/VTOS
Code: https://github.com/jinchaogjc/VTOS
License: http://creativecommons.org/licenses/by/4.0/
The gist: Vision foundation tools such as open-vocabulary detectors, segmentation models, and post-processing operators are powerful building blocks for computer vision, but their effectiveness depends heavily
Terminology
Abstract
Vision foundation tools such as open-vocabulary detectors, segmentation models, and post-processing operators are powerful building blocks for computer vision, but their effectiveness depends heavily on how they are orchestrated: which tools are used, in what order, with what parameters, and under what visual conditions. Existing visual-programming agents typically generate a fixed solution pipeline, making them brittle under dense objects, occlusion, small targets, and domain shift. We introduce VTOS (Vision Tools Orchestration Search), a framework for adaptive visual tool orchestration through joint solution-observer search. VTOS co-searches executable solution programs that compose vision tools such as Grounding DINO, SAM, NMS, and slice-and-detect, together with observer programs that diagnose candidate solutions, identify failure modes, and generate actionable feedback. These observations are accumulated in a shared VisionThoughts knowledge base to guide subsequent search. We evaluate VTOS through two case studies: dense object counting on LVIS-Count and zero-shot plant-disease segmentation on PlantSeg-OOD, which stress different orchestration challenges including threshold calibration, NMS, slicing, mask refinement, and domain generalization. Across both tasks, VTOS outperforms static tool pipelines and agentic visual-programming baselines, specifically in complex settings such as dense, occluded scenes and out-of-distribution segmentation where static pipelines leave measurable headroom, rather than in standard tasks where a single well-calibrated tool already approaches its ceiling.
Sources
- MetaPrompting: Learning to Learn Better Prompts
- AIDE: AI-Driven Exploration in the Space of Code
- What Does It Take to Be a Good AI Research Agent? Studying the Role of Ideation Diversity
- Autonomous Computer Vision Development with Agentic AI
- Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection
- System Prompt Optimization with Meta-Learning
- LVIS: A Dataset for Large Vocabulary Instance Segmentation
- REFINER: Reasoning Feedback on Intermediate Representations
- MLE-Dojo: Interactive Environments for Empowering LLM Agents in Machine Learning Engineering
- CORAL: Towards Autonomous Multi-Agent Evolution for Open-Ended Discovery
- AI Research Agents for Machine Learning: Search, Exploration, and Generalization in MLE-bench
- Learning Transferable Visual Models From Natural Language Supervision
- Learning To Count Everything
- SAM 2: Segment Anything in Images and Videos
- R&D-Agent: An LLM-Agent Framework Towards Autonomous Data Science
- AutoMMLab: Automatically Generating Deployable Models from Language Instructions for Computer Vision Tasks
- TextGrad: Automatic "Differentiation" via Text
- Synthetic Sandbox for Training Machine Learning Engineering Agents
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models