S4VY: Segment Anything in Feed-Forward 4D Visual Geometry
cs.CV
Submitted: 2026-09-29
Updated: 2026-10-03
Terminology
Sources
- Qwen3-VL Technical Report
- ARKitScenes: A Diverse Real-World Dataset For 3D Indoor Scene Understanding Using Mobile RGB-D Data
- Contrastive Lift: 3D Object Instance Segmentation by Slow-Fast Contrastive Fusion
- Matterport3D: Learning from RGB-D Data in Indoor Environments
- Grounded 3D-LLM with Referent Tokens
- Mask2Former for Video Instance Segmentation
- SpaCeFormer: Fast Proposal-Free Open-Vocabulary 3D Instance Segmentation
- SAM-V: Geometry-Aware Segment Anything for Multi-View Instance Segmentation
- GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning
- LoRA: Low-Rank Adaptation of Large Language Models
- Segment Any 4D Gaussians
- FAST3DIS: Feed-forward Anchored Scene Transformer for 3D Instance Segmentation
- 4DVLT: Dynamic Scene Understanding with Worldline-Centered Vision-Language Tracking
- IGGT: Instance-Grounded Geometry Transformer for Semantic 3D Reconstruction
- SAM3-I: Segment Anything with Instructions
- LensWalk: Agentic Video Understanding by Planning How You See in Videos
- The 2017 DAVIS Challenge on Video Object Segmentation
- SegVGGT: Joint 3D Reconstruction and Instance Segmentation from Multi-View Images
- ReScene4D: Temporally Consistent Semantic Instance Segmentation of Evolving Indoor 3D Scenes
- OpenMask3D: Open-Vocabulary 3D Instance Segmentation
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models