Robust Promptable Video Object Segmentation

arXiv:2605.12006 · cs.CV · Submitted 2026-05-12 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "Robust Promptable Video Object Segmentation".

Tom: Promptable video object segmentation (PVOS) models suffer substantial performance degradation when faced with input corruptions, which severely limits their deployment in safety-critical domains such as autonomous vehicles and robotics.

Jane: First, who's behind it and why it matters.

Title and authors: Tom: Let's talk about the title and who put this paper together. "Robust Promptable Video Object Segmentation" tells us right away the goal is to make promptable segmentation work even when the video input isn't perfect. The authors are a team from POSTECH, Google, and ETH Zurich who clearly have a strong foundation in this area of computer vision.

Jane: It’s important to remember that the problem they're solving isn't just about making things look better in noisy videos; it’s about ensuring the segmentation remains accurate and consistent across different frames despite those corruptions. That consistency is what makes it useful for tracking objects over time.

Lu: The authors are tackling a fundamental limitation that previous approaches couldn't handle well, specifically that they fail to account for how degradation affects different objects differently within the same scene, which is a key point they want to address with their new methodology.

Meng: Does this mean we're moving toward systems that don't just correct the image locally but understand the context of multiple objects simultaneously under adverse conditions? That sounds like a significant leap for practical deployment.

Lalam: I think the focus on object-specific representations is really powerful because it implies a level of internal model awareness that goes beyond simple pixel-level noise filtering.

The paper's summary: Tom: So, what’s the core idea they propose in "Robust Promptable Video Object Segmentation"? Simply put, they introduce a method called Memory-object-conditioned Gatedrank Adaptation, or MoGA, which is designed to handle object-specific degradation while keeping predictions consistent over time.

Jane: That’s a lot of technical jargon there, but I can break it down for you. The paper suggests that instead of treating the entire video frame uniformly when it's degraded, the model should remember unique characteristics for each individual object from previous frames and use those memories to adapt its predictions now.

Lu: The key insight they present is leveraging those object-specific representations that are already present in modern video segmentation models to condition how they robustify the input data, which is a sophisticated way to handle temporal consistency.

Meng: So, if we think about it practically, this means when the fog gets thicker or there's motion blur in one frame, the system doesn't just guess; it remembers what that specific car looked like before and adjusts its segmentation based on that stored memory.

Lalam: I see this as a step toward building AI systems that have a sort of short-term, object-focused episodic memory during video processing, which really helps maintain identity across sequences.

The paper's improvements: Tom: Now let's look at what they actually did to make this work better. They created a new benchmark suite that includes real-world datasets with over two thousand five hundred annotated object masks under conditions like rain, fog, and snow, alongside synthetic data generated by applying eight diverse corruption types with temporal variations.

Jane: That benchmark is substantial because it covers both real-world scenarios and controlled synthetic testing. They also explicitly highlighted the gap they were filling: existing robust image segmentation methods fail because they treat every part of the frame the same way, whereas this approach accounts for how different objects are affected differently by those adverse conditions.

Lu: The methodology introduces several steps, starting with decomposing a learnable weight matrix into rank-one components and then conditioning a gating mechanism on an object pointer maintained in a memory bank that accumulates historical characteristics specific to each object.

Meng: The integration of this into the SAM2 architecture is interesting; they are conditioning the self-attention and cross-attention layers using these object pointers, which means the adaptation is applied directly where segmentation happens.

Lalam: The training process is clever because they only train the gating modules through segmentation loss, learning to select adaptation paths based purely on how well it improves the object segmentation quality.

Conclusion: Tom: To wrap things up on "Robust Promptable Video Object Segmentation," the main conclusion is that MoGA combined with SAM2 significantly outperforms per-frame approaches across all metrics, reaching seventy-one point eight percent J andF on MVSegadv and sixty-four point five percent J andF on ACDC-Video, demonstrating that memory-based conditioning helps substantially more than frame-wise methods.

Jane: It really shows that by focusing on object identity across frames, we can achieve better results when dealing with complex real-world degradations compared to just trying to make every single frame look perfectly clear independently.

Lu: The parameter efficiency they report is also noteworthy; they manage to train MoGA+SAM2 with only one point one million parameters while achieving those strong performance numbers, which shows a very effective use of the model's capacity.

Meng: That parameter count makes a huge difference for practical deployment; you can get high performance without needing the massive computational footprint of fully fine-tuned models, which is essential for real-time systems.

Lalam: I think the implications for AI culture here are that we are moving toward building models with a more structured way of remembering things, which could lead to much more coherent and trustworthy AI agents in complex visual environments.

Tom: So, to summarize this Robust Promptable Video Object Segmentation paper, it’s about using memory-object-conditioned Gatedrank Adaptation to ensure temporal consistency in video segmentation under challenging conditions. We'll be moving on now to talk about some of the other interesting papers we've been looking at today.

Jane: That was a fantastic deep dive into MoGA and its benchmark results, Tom. It really illustrates how specific memory conditioning can solve temporal inconsistencies in vision tasks.

Lu: The work lays a solid foundation for future research into how we can generalize these object-specific memories to even more complex scenarios, perhaps across different modalities.

Meng: I'm ready to hear what the next paper is focusing on, especially since this paper shows how much structure you can find in the data itself.

Lalam: I’m looking forward to hearing how other researchers are using these memory concepts to build more reliable and context-aware AI systems.

Sohyun Lee, Yeho Gwon, Lukas Hoyer, Konrad Schindler, Christos Sakaridis, Suha Kwak

POSTECH · Google

cs.CV

Submitted: 2026-05-12

Updated: 2026-09-29

Comments: Accepted to CVPR 2026

Project page: https://sohyun-l.github.io/RobustPVOS_project_page/Abstract

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 89/100

The gist: Promptable video object segmentation (PVOS) models suffer substantial performance degradation when faced with input corruptions, which severely limits their deployment in safety-critical domains such

Key concepts

Promptable Video Object Segmentation (PVOS)
Models that use prompts for video object segmentation suffer performance drops when the input video is corrupted. This paper focuses on making these models robust so they remain accurate even with imperfect video data, which is crucial for safety-critical applications.
Memory-object-conditioned Gatedrank Adaptation (MoGA)
This proposed method introduces a technique where the model remembers unique characteristics for each individual object from previous frames. It uses these stored memories to adapt predictions in the current frame, allowing the system to handle degradation based on what it knows about that specific object.
Object-specific representations
The authors focus on leveraging existing representations of objects within video models. By conditioning how the input data is robustified using these object-specific characteristics, they achieve a level of internal model awareness beyond simple pixel noise filtering.
Temporal consistency
This refers to ensuring that the segmentation predictions for an object remain consistent across different frames in a video sequence. The paper shows that memory-based conditioning helps maintain this identity over time when facing adverse conditions like fog or motion blur.

Terminology

Summary

Promptable video object segmentation (PVOS) models suffer substantial performance degradation when faced with input corruptions, which severely limits their deployment in safety-critical domains such as autonomous vehicles and robotics. This paper addresses this critical gap by introducing RobustPVOS, a comprehensive study focused on developing robust PVOS models capable of tracking and segmenting objects across frames despite adverse conditions. The authors establish a novel benchmark encompassing real-world evaluation datasets with dense object-level annotations under natural corruptions and synthetic training data generated through temporally varying degradations. Their primary contribution is the proposal of Memory-object-conditioned Gatedrank Adaptation (MoGA), a method designed to handle object-specific degradation while ensuring temporal consistency in predictions.

RobustPVOS Benchmark Construction

The authors introduce the first comprehensive study on RobustPVOS by constructing a new benchmark suite. This benchmark is composed of two distinct parts:

  1. Real-world evaluation datasets: They curate two sets, including 351 video clips with over 2,500 annotated object masks from existing collections captured under conditions like rain, fog, snow, and nighttime scenarios. They also construct a dedicated test subset called MVSeg-adv by retaining only clips under challenging conditions such as low light, rain, snow, noise, and motion blur.

  2. Synthetic training data: They generate synthetic training data by applying eight diverse corruption types with temporal variations to annotated videos from established VOS datasets. These corruptions include color jitter, Gaussian noise, ISO noise, motion blur, resampling blur, and environmental effects like fog [4], rain [13], and snow [19].

MoGA Method Overview

The proposed method is Memory-object-conditioned Gatedrank Adaptation (MoGA). The key insight driving this approach is that object-specific representations maintained across frames, which are inherent in modern video segmentation models, naturally capture how each object is uniquely affected by degradations over time. MoGA achieves robustification by conditioning the process on these object representations stored in memory.

The mechanism involves several key steps:

  1. (Rank-1) Component Decomposition: The learnable weight matrix of the low-rank adapter, denoted as ∆W, is decomposed into R (rank-1) components, where ∆W = PR with bi ∈ R D and ai ∈ R K. This decomposition allows for a flexible construction of the weight matrix based on the input characteristics.

  2. Memory-object-conditioned Gating: The core innovation is conditioning the gating mechanism on object information stored in a memory bank. SAM2 maintains an object pointer mo ∈ O in this memory bank, which accumulates historical characteristics specific to each object. A gating module g(·) takes an object pointer as input and computes object-specific binary gating masks zo = g(mo) ∈ [0, 1] R.

  3. Differentiable Selection: To enable differentiable learning with discrete selection, they apply Gumbel-sigmoid relaxation to the logit vectors derived from the MLP applied to the object pointer: tilde z o,i = σ ((1/τ) (alpha o,i + G i)).

Training and Inference Process

MoGA is integrated into SAM2’s memory attention module, specifically into the linear projections for self-attention and cross-attention. The final output h is computed as:

h = W0 x + (1/O) sum o (Delta W o x) = W0 x + (1/O) sum o [(sum i z o,i · b i a i top) x]. Here, ∆Wo is the object-specific adapter for object o. The gating modules are trained solely through the segmentation loss: Lseg combines focal and dice losses. During training, they learn to select appropriate adaptation paths solely through the segmentation loss. At inference, the gating becomes deterministic: zo,i = I[σ(αo,i) > 0.5].

Experimental Validation and Results

Extensive experiments validate MoGA's efficacy across both synthetic and real-world datasets. On the real-world benchmark (MVSeg-adv and ACDC-Video), MoGA+SAM2 achieved significant gains: "MoGA combined with SAM2 (MoGA+SAM2) achieves significant gains across all metrics, reaching 71.8% J &F on MVSegadv and 64.5% J &F on ACDC-Video. This demonstrates that memory-based conditioning substantially outperforms per-frame approaches. Furthermore, the method proves parameter efficient; MoGA+SAM2 trains only 1.1M parameters yet achieves 71.8% J &F," compared to a fully fine-tuned SAM2 which requires 80.9M parameters and more memory.

Improvements for AI systems

Here are the specific improvements and capabilities derived from the RobustPVOS (MoGA) method:

  1. Improve robustness of promptable video object segmentation (PVOS) models against input corruptions (noise, blur, low illumination, adverse weather).

  2. Achieve temporally consistent object tracking and segmentation across frames in videos degraded by diverse temporal and spatial patterns of corruption.

  3. Enable the system to maintain high performance on clear videos while significantly mitigating performance degradation under real-world adverse conditions compared to state-of-the-art frame-wise restoration or independent robust image segmentation methods.

  4. Enhance parameter efficiency: achieve robust PVOS performance comparable to or exceeding fully fine-tuned models with significantly fewer trainable parameters (MoGA+SAM2 trains 1.1M parameters vs. SAM2 fine-tuned at 80.9M).

  5. Provide object-specific adaptation during inference by using object representations maintained in a memory bank to condition the robustification process, ensuring that different tracked objects are handled uniquely yet coherently over time.

The improved AI system (MoGA+SAM2) can perform the following specific tasks:

  1. Accurately segment and track arbitrary objects in video sequences given free-form prompts, even when the video frames are affected by complex real-world degradations such as rain, fog, snow, motion blur, and low illumination.

  2. Generate dense object-level annotations with consistent instance identities across all frames of a sequence captured under challenging natural conditions (e.g., ACDC and MVSeg datasets).

  3. Maintain stable tracking performance on long video sequences (up to 42 seconds) where traditional models like SAM2 suffer significant performance drops, by leveraging the memory bank for temporal context.

  4. Provide reliable segmentation outputs in safety-critical domains (autonomous vehicles, robotics) where adverse environmental conditions are inevitable and require robust, temporally consistent perception.

Sources

Related papers