Seeing as Humans Do: Learning from Motion to Segment Anything Without Supervision
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Seeing as Humans Do".
Jane: The Segment Anything Model (SAM) currently relies heavily on massive manual annotations, creating a fundamental bottleneck for model scaling.
Tom: First, who's behind it and why it matters.
Paper summary: Tom: So, to recap that MoSA is this three-stage framework designed to learn a transferable objectness prior from unlabeled videos so you can get segmentation performance similar to SAM without any manual annotations. Jane, what's the main claim they are making about their approach?
Jane: The paper claims they overcome the bottleneck of needing massive manual annotations by introducing MoSA, which learns a generalized concept of objects from motion cues in videos. They achieve this by generating motion pseudo-labels first, then training a Perceptual Grouping Model to internalize that object concept, and finally transferring that knowledge into a prompt-guided architecture for segmentation on images.
Lu: I find the progression of stages really interesting; moving from generating those motion pseudo-labels in the first stage to training the PGM in the second stage, and then adapting it in the third stage shows a very thoughtful pipeline design.
Meng: The emphasis on learning a concept independent of explicit object count or categories through contrastive learning sounds like a clever way to create something truly generalizable, which is what I need for robust engineering.
Lalam: If this works, it means we could potentially train foundational models that understand the world structure implicitly just by observing how things move in massive datasets, which would drastically improve our ability to build context-aware applications.
Conclusion: Tom: So, we’ve talked about how MoSA aims to replace manual labeling with motion learning, and now we’re looking at the bigger picture with the authors and the implications of "Seeing as Humans Do: Learning from Motion to Segment Anything Without Supervision." Jane, what's your take on what this paper means for how we think about segmentation?
Jane: The title suggests that by focusing on motion, we are trying to emulate a more holistic understanding of scenes, moving beyond just static recognition. The authors are showing that learning object concepts from motion is a viable path toward scaling models without relying solely on the intensive manual annotation process.
Lu: I think the implication is that we can build systems that aren't tethered to specific training datasets for every single object type; they could learn general rules about what constitutes an 'object' based on physical interactions observed in video.
Meng: For practical deployment, this suggests a future where we can deploy vision models much faster because the pre-training phase relies on existing video data rather than waiting for perfectly curated, expensive labeled datasets for every new use case.
Lalam: This could fundamentally change how we approach AI development; instead of labeling thousands of images for every new segmentation task, we might be able to train a model once on large video collections and then adapt it with minimal supervision for specific tasks.
Tom: It really boils down to making the perception process more data-efficient by leveraging the rich information already present in raw video data. That's what this paper is pointing toward as a way forward.
Weijian Jian, *, *Xiaoyue Zhang*, *Bin Xiao*, *Chunyu Xie*, *Yixiao He*, *Yutao Liu*, Dawei Leng, Yuhui Yin
360 AI Research · University of Ottawa · Beijing University of Posts and Telecommunications
cs.CV
Submitted: 2026-09-30
Updated: 2026-09-30
Code: https://github.com/360CVGroup/MoSA
Importance score: 90/100
The gist: The Segment Anything Model (SAM) currently relies heavily on massive manual annotations, creating a fundamental bottleneck for model scaling.
Key concepts
- Multi-Granularity Motion Segmentation (MGMS)
- This pipeline uses optical flow to generate motion pseudo-labels from videos. It has two parts: one module segments rigid objects, and another handles complex moving foregrounds. Quality scores filter these masks for high accuracy.
- Perceptual Grouping Model (PGM)
- Built on a Vision Transformer, this model learns a generalized concept of 'objectness' from motion data. It uses contrastive learning to group patches belonging to the same object without needing explicit labels, capturing appearance-driven object features.
- Perceptual Grouping Contrastive Learning (PGCL)
- This specific training method contrasts positive pairs (patches in the same mask) against negative ones. It models the probability of two patches being from the same object using a learnable temperature to refine how similar objects are perceived.
Terminology
Summary
The Segment Anything Model (SAM) currently relies heavily on massive manual annotations, creating a fundamental bottleneck for model scaling. This paper introduces Motion-Grounded Segment Anything (MoSA), a highly scalable unsupervised framework that learns a transferable objectness prior from unlabeled videos to achieve segmentation performance comparable to the fully supervised SAM without requiring manual annotations.
How it works
The MoSA framework operates in three progressive stages:
-
Automatically generating multi-granularity motion pseudo-labels from large-scale video data via a Multi-Granularity Motion Segmentation (MGMS) pipeline.
-
Training a Perceptual Grouping Model (PGM) via contrastive learning to internalize a generalized, appearance-driven concept of objects from these motion pseudo-labels.
-
Transferring this learned prior into a prompt-guided architecture for segment-anything style inference on images, enabling strong zero-shot performance.
Multi-Granularity Motion Segmentation (MGMS)
The MGMS pipeline is designed to generate high-quality motion pseudo-labels from unlabeled videos using bidirectional optical flow. It comprises two complementary modules:
-
Rigid Object Segmentation Module (ROSM): This module adapts the SOLOv2 architecture to segment individual rigidly moving objects, leveraging parallax for static foreground objects and discerning object parts.
-
Motion Foreground Segmentation Module (MFSM): This module employs a U-Net architecture to delineate the entire moving foreground from background, capturing complex non-rigid deformations.
To ensure high fidelity, candidate masks are filtered using a Mask Quality Score (Squality), which combines two components: Maskness (Smaskness)—the mean prediction value for pixels exceeding a confidence threshold—and Boundary Sharpness (Ssharpness)—the proportion of boundary pixels with a gradient magnitude above a threshold. Masks with an Squality below 0.85 are discarded, and Non-Maximum Suppression (NMS) is applied to yield the final set of pseudo-labels.
Perceptual Grouping Model (PGM)
The PGM is built upon a Vision Transformer (ViT) backbone and incorporates K parallel linear projection heads to capture objects at various scales. A novel contrastive training strategy, Perceptual Grouping Contrastive Learning (PGCL), is employed to learn a generalizable notion of “objectness” independent of explicit object count or categories.
The PGCL loss contrasts positive pairs (patches within the same mask) against negative ones by modeling the probability that a patch belongs to the same object as an anchor patch, using a sigmoid function with a learnable temperature τk:
p t,j k = σ (τ k · s t k), where s t k is the cosine similarity between anchor and patch features from head k.
Adaptation to Segment Anything
The learned objectness prior from the PGM is transferred into two specialized models for inference:
-
Whole-Image Segmentation Model: This model employs a DINO pre-trained ResNet-50 backbone with a Mask2Former decoder, configured with 1000 learnable object queries, to output automatic segmentation masks for high-resolution images.
-
Promptable Segmentation Model: This model is based on the Semantic-SAM framework and uses a SwinTransformer Tiny backbone to generate hierarchical masks from user prompts, simulating user clicks during training by sampling points within the foreground of each mask as positive point prompts.
Evaluation and Results
MoSA achieves state-of-the-art performance across seven challenging benchmarks (e.g., COCO and ADE20K). For whole-image segmentation, MoSA achieves an average Average Recall (AR) of 42.1% on the SA-1B dataset, significantly outperforming existing unsupervised methods like UnSAM [38]. For point-based promptable segmentation on COCO Val2017, MoSA achieves a MaxIoU of 41.6% and an OracleIoU of 63.4%, surpassing the previous SOTA, UnSAM [38], by +1.3% and +3.9% respectively. Ablation studies confirm that the combination of ROSM and MFSM is essential, as their individual contributions are complementary for capturing multi-scale motion cues, leading to a final score of 36.6% AR1000 when both are included in the MGMS pipeline. The study concludes that learning general object perception from large-scale unlabeled motion is a feasible and scalable alternative to annotation-driven pipelines.
Limitations
A practical failure mode identified is degraded motion estimation in lowlight environments, such as night scenes, where unreliable optical flow can weaken supervision for fine structures. Future research will focus on improving robustness to these flow-degraded conditions. The current work focuses on extracting general segmentation knowledge from motion cues with the goal of enabling zero-shot segmentation on static images without relying on explicit motion information during inference.
Improvements for AI systems
Here are the specific improvements that can be made to AI systems by implementing Motion-Grounded Segment Anything (MoSA), along with a description of what these improved systems will be capable of:
The implementation of MoSA enables the creation of a foundational, scalable segmentation system capable of achieving state-of-the-art performance comparable to fully supervised models like SAM, but trained entirely without manual annotations.
Specific improvements and capabilities include:
- textbf>Zero-Shot Multi-Granularity Segmentation from Unlabeled Video Data (Core Capability):
Based on the three-stage MoSA framework (MGMS pipeline, PGM contrastive learning, and Adaptation), the improved system can perform segmentation on completely unseen static images in a zero-shot manner. It learns a generalized concept of objectness
by observing motion patterns in massive, unlabeled videos (Kinetics-700, BDD100K) rather than relying on specific object labels.
- textbf>Generalization to Static and Weakly Moving Entities (Addressing Unsupervised Bias):
The Perceptual Grouping Model (PGM), trained via Perceptual Grouping Contrastive Learning (PGCL), overcomes the limitation of motion-only models by internalizing appearance-driven concepts. This allows the system to successfully segment objects that are completely static or exhibit only slight motion (e.g., roof tiles, stationary furniture) where traditional motion-based methods fail, significantly expanding its object vocabulary beyond moving entities.
- textbf>Handling Complex Multi-Granularity Objects (Enhanced Structural Understanding):
The Multi-Granularity Motion Segmentation (MGMS) pipeline uses two complementary modules: the Rigid Object Segmentation Module (ROSM) and the Motion Foreground Segmentation Module (MFSM). This allows the system to simultaneously capture:
-
Precise, instance-level masks of rigidly moving objects and their articulated parts.
-
Coarse, encompassing binary masks for complex, non-rigid entities like pedestrians or animals.
This duality ensures a comprehensive understanding of objects across different scales and structural complexities within a single video frame.
- textbf>Promptable Interaction with High-Resolution Fidelity (Interactive Segmentation):
By transferring the learned object priors into prompt-guided architectures (like Mask2Former for whole-image segmentation and Semantic-SAM for promptable segmentation), the system gains interactive capabilities. Users can interact with the model using point prompts, leading to high-quality, fine-grained masks across multiple granularity levels (e.g., segmenting a person's arm vs. their entire torso) directly from high-resolution images.
- textbf>Scalability and Efficiency (Reduced Annotation Bottleneck):
The system drastically reduces the dependence on costly manual annotations (SA-1B). The MoSA framework is shown to achieve performance comparable to SAM while requiring significantly less training data (1% of SA-1B) and fewer model parameters, making it a highly scalable alternative for deploying general-purpose segmentation models in resource-constrained environments.
- textbf>Robustness via Multi-Stage Filtering (High Confidence Outputs):
The inclusion of a rigorous Mask Quality Score (Squality), which combines maskness and boundary sharpness metrics, ensures that the system produces outputs with high pixel confidence and clear boundaries, minimizing segmentation noise—a critical improvement over methods relying solely on raw pseudo-labels.
Sources
- YouTube-8M: A Large-Scale Video Classification Benchmark
- A Short Note on the Kinetics-700 Human Action Dataset
- DINOv2: Learning Robust Visual Features without Supervision
- DINOv3
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models