OPERA: A Unified Omnimodal Progressive Spatio-Temporal Reasoning Agent for Referring Video Segmentation
cs.CV
Submitted: 2026-09-27
Updated: 2026-09-27
Terminology
Sources
- Ming-Flash-Omni: A Sparse, Unified Architecture for Multimodal Perception and Generation
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- SimSD: Simple Speculative Decoding in Diffusion Language Models
- Cognitive Kernel-Pro: A Framework for Deep Research Agents and Agent Foundation Models Training
- Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos
- Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement
- What Makes Low-Bit Quantization-Aware Training Work for Reasoning LLMs? A Systematic Study
- Cascaded Self-Evaluation Augmented Training for Lightweight Multimodal LLMs
- Natural-Language Agent Harnesses
- SAM 2: Segment Anything in Images and Videos
- Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks
- DeskCraft: Benchmarking Desktop Agents on Professional Workflows and Human-in-the-Loop Collaboration
- VG-Refiner: Towards Tool-Refined Referring Grounded Reasoning via Agentic Reinforcement Learning
- MC-CoT: A Modular Collaborative CoT Framework for Zero-shot Medical-VQA with LLM and MLLM Integration
- Qwen3-Omni Technical Report
- Qwen2.5-Omni Technical Report
- AdaMem: Adaptive User-Centric Memory for Long-Horizon Dialogue Agents
- R1-Omni: Explainable Omni-Multimodal Emotion Recognition with Reinforcement Learning
- ViLLa: Video Reasoning Segmentation with Large Language Model
- LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models