Vision Harnessing Agent for Open Ad-hoc Segmentation
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Vision Harnessing Agent for Open Ad-hoc Segmentation".
Jane: The Vision Harnessing Agent for Open Ad-hoc Segmentation (VASA) introduces a novel,
Tom: First, who's behind it and why it matters.
Title and authors: Tom: So we’ve talked about what VASA is, but let’s really dig into the core mechanics of how this Vision Harnessing Agent for Open Ad-hoc Segmentation actually works under the hood.
Jane: We can focus on how it moves beyond traditional methods by looking at its iterative construction process and its state management features.
Lu: The summary points out that VASA is fundamentally a training-free approach, meaning it doesn't require retraining the agent for every new type of segmentation task.
Meng: That’s important for practical deployment because continuous retraining is often too slow for rapidly evolving visual tasks in a startup environment like mine.
Lalam: The core mechanism involves formulating segmentation as an iterative visual construction process over a persistent working mask, which is carried across interaction rounds.
Tom: And the agent plans these operations by selecting tools like ADD, REMOVE, or REPLACE based on candidate masks it gets back from the segmentation model.
Jane: It seems like the power comes from this loop where it doesn't just output a final answer but instead reasons through a sequence of visual edits until the goal is met.
Lu: The state management is crucial because "found, removed, or refined regions remain available for subsequent reasoning," which keeps the agent from starting over every time.
Meng: That persistent memory of what has been done must be very efficient to manage without slowing down the entire process.
Lalam: The paper stresses that this planning is guided by tracking interaction history and budget, which lets the agent revise its strategy when intermediate results don't satisfy the query.
Tom: So it’s not just one big prediction step; it’s a series of smaller, verifiable visual steps that lead to the final output, which is a key distinction from prompt refinement agents.
Jane: That makes sense; instead of trying to write the perfect sentence upfront, you have an agent that experiments visually and learns by doing.
Lu: This level of structured planning moves the system toward a more goal-directed reasoning process for visual tasks.
Meng: I’m curious if this iterative editing process adds too much latency compared to a single, optimized inference pass, but the results suggest it’s worth that trade-off.
Lalam: The paper demonstrates this approach successfully on PARS and RefCOCOm, where VASA outperforms SAM3 Agent by fourteen to twenty-five percent on the ad-hoc split.
The paper's summary: Tom: Now that we understand the mechanics, let’s focus on what specific improvements this approach brings to the existing methods compared to what we see in SAM3 Agent or similar agentic baselines.
Jane: The primary improvement seems to be moving from simple text-to-mask retrieval to a genuine vision harnessing agent approach that employs persistent visual state.
Lu: It’s about giving the AI the ability to perform open ad-hoc segmentation by constructing complex concepts on the fly using a structured workflow.
Meng: Can you elaborate on what that "structured workflow" means in terms of complexity? Is it just a simple if-then structure, or something more dynamic?
Lalam: It’s more than a simple structure; it involves programmatically pixel-level Boolean updates executed by a deterministic program to apply the selected edit operation to the working mask in place.
Tom: That level of programmatic control over the mask editing is what really sets it apart from just letting a model try different text prompts repeatedly.
Jane: It means the agent can handle complex constraints like exclusions and compositions by iteratively segmenting and then removing specific regions as needed.
Lu: The system enhances robustness through explicit workflow rules and visual scrutiny over intermediate masks, allowing it to detect missing regions or oversegmentation.
Meng: That self-correction mechanism sounds like a solid way to handle the inherent uncertainty in complex visual tasks without needing a massive safety layer.
Lalam: Furthermore, they include recovery mechanisms for failures like stalled progress or response-formatting errors, letting the agent retry with a different prompt while preserving its history.
Tom: So in essence, the improvement is making the agent proactive—it plans a strategy rather than just reacting to text revisions.
Jane: It really demonstrates how programming agents with "task knowledge, VLM behavior, visual routines" allows them to handle open-world environments better.
Lu: The impact is that we are pointing toward AI agents capable of complex visual construction because they aren't limited to just retrieving pre-learned concepts anymore.
The paper's improvements: Tom: We’ve covered a lot about the Vision Harnessing Agent for Open Ad-hoc Segmentation, and it seems like the main message is that we are moving toward agents that can build visual concepts dynamically.
Jane: It really boils down to this: instead of just asking for a known object, we give the AI a persistent visual state and a structured way to reason through edits until the desired mask is constructed.
Lu: The implication for the field is that we are defining what it means to build complex segmentation tasks in an open-ended setting, moving past simple retrieval models.
Meng: From a practical view, this suggests a future where agents could localize task-relevant concepts defined by context or user intent in environments that change constantly.
Lalam: I see this as a significant step toward more capable AI systems by programming them with these visual routines and failure-aware workflows.
Tom: So, to wrap up, the VASA framework introduces a way to handle open ad-hoc segmentation by coupling a VLM agent with a segmentation model through this persistent working mask and iterative visual construction process.
Jane: We’ve seen how it plans operations, checks its work visually, and recovers from errors to achieve much higher performance on benchmarks like PARS compared to SAM3 Agent.
Lu: It’s a strong contribution because it establishes a new way for AI agents to tackle the complexity of constructing visual concepts from scratch.
Meng: As an engineer, I think the focus on structured tool calls and deterministic execution is what makes this approach more reliable for production systems.
Lalam: Overall, this paper on Vision Harnessing Agent for Open Ad-hoc Segmentation shows a clear path for AI agents by programming them with visual routines and failure-aware workflows.
Conclusion: Tom: So, to wrap up what we've discussed about the Vision Harnessing Agent for Open Ad-hoc Segmentation, we’ve seen how this new framework lets AI agents construct complex visual ideas on the fly using persistent visual states and iterative reasoning.
Jane: It’s really impressive how VASA shifts the focus from just refining a text prompt to programming an agent with actual "visual routines" so it can handle open-world tasks.
Lu: The theoretical potential here is huge; thinking about agents that can program their own visual logic based on constraints feels like we're getting closer to truly general visual reasoning.
Meng: From an engineering standpoint, the ability for the agent to manage its working mask persistently across rounds is a big plus for building reliable applications that need to handle changing visual requirements.
Lalam: I think this advancement in structured construction capability will improve how we train and deploy agents, because having a verifiable, step-by-step visual construction process makes the entire system much more trustworthy.
Tom: Exactly! The results showing VASA consistently outperforming SAM3 Agent by those significant margins on PARS really show that this iterative, stateful approach is better at handling those ad-hoc splits.
Jane: It’s clear that the performance gains come from that structured workflow and the rigorous scrutiny the agent applies to its intermediate masks, which helps it avoid those common pitfalls like oversegmentation <ref:two thousand six hundred five point one nine four one zero#pg2.
Lu: And I think we should really look at how this concept of "programming visual routines" connects with other areas, perhaps even in how we model dynamic interactions in physical simulations, which is something I've been thinking about lately <ref:two thousand six hundred five point one nine four one zero#pg1.
Meng: That idea of context-aware reasoning that doesn't just rely on static prompts is exactly what we need for autonomous systems operating outside of a controlled lab setting <ref:two thousand six hundred five point one nine four one zero#pg0.
Lalam: For me, the most impactful aspect is how this framework allows us to build agents that can adapt their internal strategies based on visual feedback, which could fundamentally change how we design AI cultural tools <ref:two thousand six hundred five point one nine four one zero#pg2.
Tom: So, to recap, the Vision Harnessing Agent for Open Ad-hoc Segmentation gives us a powerful way to build complex concepts by programming agents with persistent state and iterative visual construction.
Jane: It’s a really exciting piece of work that shows how detailed, long-form descriptions can actually guide an agent through the necessary steps to create what it needs <ref:two thousand six hundred five point one nine four one zero#pg2.
Lu: This paper lays some interesting groundwork for future work where agents handle more abstract, relational concepts rather than just pixel masks <ref:two thousand six hundred five point one nine four one zero#pg1.
Meng: We’ll keep an eye on how this iterative planning translates into real-world latency figures, because that’s the main hurdle we need to clear for deployment <ref:two thousand six hundred five point one nine four one zero#pg2.
Lalam: I'm really optimistic about the future of this approach in creating more sophisticated and adaptive AI systems, and I look forward to seeing how these visual routines evolve <ref:two thousand six hundred five point one nine four one zero#pg2.
Zilin Wang, Stella X. Yu
University of Michigan
cs.CV
Submitted: 2026-05-19
Updated: 2026-09-29
Code: https://github.com/Wayne2Wang/VASA
Importance score: 81/100
The gist: The Vision Harnessing Agent for Open Ad-hoc Segmentation (VASA) introduces a novel, training-free framework that enables AI agents to construct visual concepts on the fly by iteratively reasoning
Key concepts
- Training-free approach
- VASA is fundamentally a training-free method for segmentation. This means the agent does not need to be retrained for every new type of segmentation task, which is important for fast deployment in evolving visual environments.
- Iterative visual construction process
- Segmentation is framed as an iterative process where the agent works over a persistent working mask. It plans operations by selecting tools like ADD, REMOVE, or REPLACE based on masks returned from the segmentation model, reasoning through a sequence of visual edits instead of just one prediction.
- Persistent working mask
- This is a mask that is carried across interaction rounds. The agent keeps track of what has been found, removed, or refined in this persistent state. This memory allows the agent to reason through subsequent steps without starting over every time.
- Programmatic pixel-level Boolean updates
- The planning involves executing deterministic programs to apply selected edit operations directly to the working mask in place. This provides programmatic control over mask editing, allowing the agent to handle complex constraints like exclusions and compositions.
Terminology
Summary
The Vision Harnessing Agent for Open Ad-hoc Segmentation (VASA) introduces a novel, training-free framework that enables AI agents to construct visual concepts on the fly by iteratively reasoning over persistent visual states. This work addresses the challenge of open ad-hoc segmentation—where concepts are defined by parts, relations, exclusions, or collections rather than existing learned masks—by coupling a Vision-Language Model (VLM) agent with a segmentation foundation model through a structured workflow. VASA’s significance lies in moving beyond simple prompt refinement by programming agents with task knowledge, VLM behavior, visual routines, working memory, and failure-aware workflows,
pointing toward AI agents capable of complex visual construction in open-world environments.
The Core Problem: Open Ad-hoc Segmentation
Segmentation becomes difficult when the concept is open and ad-hoc; unlike known concepts (e.g., cat
), an ad-hoc concept might require constructing a mask from parts, relations, exclusions, or collections.
This setting extends open ad-hoc categorization by imposing a stricter demand: the grounding must be at the pixel level. The paper illustrates this gap by comparing methods on prompts like the cat’s head without the ears and eyes,
showing that existing agentic baselines like SAM3 Agent fail because they treat each attempt as a fresh retrieval, whereas VASA builds visual progress persistently.
VASA's Vision Harness Engineering Workflow
VASA is designed as a training-free framework that formulates segmentation as an iterative visual construction process over a persistent working mask.
It couples a VLM agent with SAM3 through a workflow that coordinates planning, state management, tool invocation, action constraints, and error recovery. The key components of this workflow are:
-
A persistent working mask initialized as empty and carried across interaction rounds to record what visual construction has already been achieved.
-
Structured tool calls where the VLM agent selects an operation (ADD, REMOVE, or REPLACE) based on candidate masks returned by the segmentation model (SAM3).
-
Programmatic pixel-level Boolean updates executed by a deterministic program to apply the selected edit operation to the working mask in place.
Long-Horizon Planning and State Management
Unlike prompt-refinement agents that restart from each revised text prompt, VASA plans over multiple tool-use steps rather than directly predicting the final mask. The VLM agent analyzes the target concept's structure and selects strategies such as “direct retrieval,” “undersegment-and-add,” or “oversegment-and-remove.” This planning is guided by tracking interaction history and budget, allowing the agent to revise its strategy when intermediate results do not satisfy the query. State management is crucial, as Found, removed, or refined regions remain available for subsequent reasoning.
Scrutiny and Error Recovery
VASA constrains the VLM agent through explicit workflow rules and visual scrutiny over intermediate masks. After each round, the agent checks whether the working mask satisfies inclusion, exclusion, structural, and relational constraints expressed in the query. This scrutiny allows VASA to detect missing regions, oversegmentation, concept confusion, or failed exclusions.
To ensure stable long-horizon interaction, it includes recovery for failures such as stalled progress and response-formatting errors,
allowing the agent to retry with a different prompt while preserving history.
Evaluation and Results
VASA is evaluated on two benchmarks: PARS (PartImageNet Ad-hoc Referring Segmentation), which uses detailed long-form queries, and RefCOCOm, a standard multi-granularity referring segmentation benchmark. On PARS, VASA consistently outperforms SAM3 Agent by 13.5%–25.3% across metrics like gIoU and cIoU on the ad-hoc split, demonstrating superior ability to handle complex visual constructions. Furthermore, qualitative results show that VASA more faithfully isolates requested regions in both PARS and RefCOCOm compared to baselines, confirming that long-form descriptions are not merely longer prompts but provide structured guidance for decomposing and constructing open, ad-hoc visual concepts when coupled with an appropriate vision harness.
Contributions
The paper introduces three main contributions: 1) the definition of open ad-hoc segmentation; 2) VASA, the first vision harnessing agent for this setting, featuring persistent visual state and long-horizon construction; and 3) PARS, a new benchmark that uses detailed long queries to evaluate open ad-hoc segmentation. The work suggests a path for AI agents by programming them with visual routines
and failure-aware workflows.
Limitations
The primary limitations identified are inference latency due to iterative reasoning, dependence on the underlying foundation models (if the segmenter cannot localize primitives), and the current scope of PARS being derived from part-level concepts in PartImageNet.
Improvements for AI systems
Here are specific improvements to AI systems based on the VASA framework described in the paper:
-
The core improvement is shifting from simple text-to-mask retrieval (like SAM3 Agent) to a
vision harnessing agent
approach that employs persistent visual state and iterative, executable visual construction. -
The improved system can perform open ad-hoc segmentation by constructing complex concepts on the fly using a structured workflow:
Discussing the capability: The system can segment arbitrary, task-contingent visual concepts that do not exist as single learned masks (e.g., the cat’s right paw and the stick she is reaching for,
or the cat head without ears and eyes
).
- The system improves reasoning by replacing shallow text prompt refinement with a multi-round, stateful visual construction process:
Discussing the capability: The agent can execute a sequence of visual operations (ADD, REMOVE, REPLACE) based on intermediate mask inspection. This allows it to handle complex constraints like exclusions and compositions (e.g., iteratively segment the cat head and then remove the ears).
- The system enhances robustness through error recovery and scrutiny mechanisms:
Discussing the capability: The agent can detect failures (stalled progress, incorrect exclusions) via visual scrutiny of intermediate masks and recover by revising its strategy or prompt, ensuring long-horizon reasoning is stable.
- The system enables high-precision localization for fine-grained tasks:
Discussing the capability: By leveraging detailed long-form queries (like those in PARS), the agent can achieve superior segmentation performance, especially at the part and object level, reducing cross-concept confusion (measured by xIoU).
- The system generalizes beyond specific benchmarks to real-world applications:
Discussing the capability: The framework can be applied to robotics and embodied AI where agents must localize task-relevant concepts defined by context, function, or user intent in open-world environments.
In summary, the improved AI system moves from a retrieval tool to a sophisticated agent capable of programming
segmentation tasks by giving it visual routines and working memory.
Sources
- Segment Anything
- SAM 2: Segment Anything in Images and Videos
- Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
- Qwen2.5-VL Technical Report
- Qwen3-VL Technical Report
- Improved Baselines with Visual Instruction Tuning
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- LISA: Reasoning Segmentation via Large Language Model
- Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement
- PartImageNet: A Large, High-Quality Dataset of Parts
- Learning Transferable Visual Models From Natural Language Supervision
- Sigmoid Loss for Language Image Pre-Training
- SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision
- Extract Free Dense Labels from CLIP
- SCLIP: Rethinking Self-Attention for Dense Vision-Language Inference
- High-Quality Mask Tuning Matters for Open-Vocabulary Segmentation
- GroupViT: Semantic Segmentation Emerges from Text Supervision
- Open Vocabulary Semantic Segmentation with Patch Aligned Contrastive Learning
- Scaling Open-Vocabulary Image Segmentation with Image-Level Labels
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models