Learning to Reason with Persistent Object States for Video Instance Segmentation
cs.CV
Submitted: 2026-09-28
Updated: 2026-09-28
Terminology
Sources
- SAM 3: Segment Anything with Concepts
- Mask2Former for Video Instance Segmentation
- XMem: Long-Term Video Object Segmentation with an Atkinson-Shiffrin Memory Model
- Mind the Gap: Disentangling Performance Bottlenecks in Video Instance Segmentation
- Latest Object Memory Management for Temporally Consistent Video Instance Segmentation
- 4DVLT: Dynamic Scene Understanding with Worldline-Centered Vision-Language Tracking
- VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System
- TSA: Temporal Slot Activation for Persistent Object-Centric Video Representation
- SAM3-DMS: Decoupled Memory Selection for Multi-target Video Segmentation of SAM3
- Toward Open-World Video Segmentation over Long Horizons
- Dual-State Slot Attention: Decoupling Appearance and Identity for Video Object-Centric Learning
- Accelerating Streaming Video Large Language Models via Hierarchical Token Compression
- AI for Service: Proactive Assistance with AI Glasses
- EvoStreaming: Your Offline Video Model Is a Natively Streaming Assistant
- MoSAM: Motion-Guided Segment Anything Model with Spatial-Temporal Memory Selection
- DVIS++: Improved Decoupled Framework for Universal Video Segmentation
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models