Learning from Multimodal Pseudo-Labels for Robust Open-Vocabulary Instance and Panoptic Segmentation

arXiv:2608.11681 · cs.CV, cs.AI, cs.MM · Submitted 2026-08-12 · Read on arXiv

Duy Tran Thanh, Yeejin Lee, Byeongkeun Kang

Seoul National University of Science and Technology · Chung-Ang University

cs.CV, cs.AI, cs.MM

Submitted: 2026-08-12

Updated: 2026-08-13

Comments: 14 pages

Journal ref: Neurocomputing, 2026

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 75/100

The gist: This paper addresses the challenges of open-vocabulary instance segmentation (OVIS) and open-set panoptic segmentation (OSPS), which aim to recognize both predefined and unseen object categories

Terminology

Summary

This paper addresses the challenges of open-vocabulary instance segmentation (OVIS) and open-set panoptic segmentation (OSPS), which aim to recognize both predefined and unseen object categories without exhaustive human annotations. Existing methods often suffer from noisy pseudo-masks, limited visual-textual grounding, and difficulty handling synonyms or out-of-vocabulary (OOV) words. To overcome these challenges, the authors propose a multimodal framework that leverages pre-trained vision-language models for automatic pseudo-label generation, CLIP-guided synonym filtering, and GPT-based caption reconstruction. In their target-vocabulary-assisted pseudo-labeling setting, the framework first constructs pseudo segmentation masks, descriptive captions, and semantically aligned synonym sets using Grounded SAM, LLaVA, and CLIP, providing multimodal supervision without manual annotation. They then enhance visual-textual alignment through three complementary training objectives: an extended grounding loss that incorporates visually grounded synonyms, a semantic consistency loss, and a generative caption reconstruction loss. Extensive experiments on the COCO dataset demonstrate that the proposed method consistently outperforms previous state-of-the-art approaches under this protocol, achieving substantial improvements on both OVIS and OSPS benchmarks.

The main contributions are summarized as follows:

  • "We introduce an automated multimodal pipeline that leverages pre-trained vision-language models to generate pseudo segmentation masks, descriptive captions, and semantically aligned synonym sets, providing additional multimodal supervision without manual annotations under a target-vocabulary-assisted pseudo-labeling protocol."

  • We propose a semantic consistency loss and an extended grounding loss that leverage both predefined category names and their visually grounded synonyms to improve generalization and robustness to vocabulary variations.

  • We introduce a GPT-based generative caption reconstruction loss that enhances visual-textual reasoning by reconstructing masked captions conditioned on visual features.

  • We demonstrate that the proposed method consistently outperforms previous state-of-the-art approaches on both OVIS and OSPS benchmarks using the COCO dataset under the target-vocabulary-assisted evaluation protocol.

The method adopts a Mask2Former-based segmenter as the base architecture. During training, the model uses labeled base-class data and pseudo-labels generated for novel classes. The pseudo-label generation pipeline uses Grounded SAM (combining Grounding DINO and SAM) to produce pseudo masks for novel categories, LLaVA to generate descriptive pseudo captions, and CLIP-based multimodal filtering to select visually grounded synonyms. The training objectives include the standard classification and mask losses, an extended grounding loss (Lgr) that aligns visual embeddings with both novel-class names and synonyms, a semantic consistency loss (Lcons) that enforces consistency between synonymous word embeddings, and a generative caption reconstruction loss (Lrecon) that reconstructs masked caption tokens conditioned on visual features. The total loss is a weighted sum of these components.

Quantitative results for OVIS show that compared with the original CGG method, the proposed method improves novel-class AP by 22.1 and 22.0 points under the constrained and generalized settings, respectively. For base classes, improvements are 1.0 and 1.4 AP points. Compared with a controlled baseline (CGG†, which uses the same Grounded-SAM pseudo masks but original CGG losses), the proposed method improves novel-class AP by 6.5 and 6.8 points in the constrained and generalized settings. For OSPS, the method achieves absolute improvements of 18.0, 11.5, and 7.3 PQ points for unknown classes under the 20%, 10%, and 5% unknown settings, respectively, compared with the previous state-of-the-art. Known-class performance slightly decreases, indicating a trade-off toward better unknown-class recognition.

Ablation studies show that replacing the original CGG grounding loss with the extended grounding loss improves novel AP from 45.1 to 49.5 (constrained) and from 43.6 to 47.8 (generalized). Adding the semantic consistency loss further improves results to 50.9 and 49.3, and adding the generative caption reconstruction loss yields the best performance of 51.6 and 50.4. The CLIP-based synonym filtering is also shown to be effective: without filtering, novel AP is 47.1 and 44.0, while with filtering it improves to 51.6 and 50.4. The inference-time computational overhead is modest, with parameter increases from 35.6M to 38.4M (7.9%) and GFLOPs from 227.5 to 232.9 (2.4%). The auxiliary pseudo-label generation modules are used only during training and are removed at inference.

Improvements for AI systems

Improvements to AI Systems:

  1. Multimodal Pseudo-Label Generation Pipeline
  • Integrate a pipeline combining Grounded SAM, LLaVA, and CLIP to automatically generate pseudo segmentation masks, descriptive captions, and synonym sets for unseen categories without manual annotation.

  • This enables AI systems to learn open-vocabulary recognition from unlabeled data, reducing reliance on exhaustive human labeling.

  1. Extended Grounding Loss with Visually Grounded Synonyms
  • Modify the grounding loss to align visual embeddings not only with predefined category names but also with CLIP-filtered synonyms that are visually grounded in the image.

  • This improves robustness to vocabulary variations and handles synonyms or out-of-vocabulary words more effectively, leading to better generalization to novel classes.

  1. Semantic Consistency Loss for Synonym Embeddings
  • Enforce consistency between word embeddings of synonymous terms (e.g., puppy and dog) during training.

  • This reduces semantic drift and improves the model’s ability to recognize the same object across different linguistic expressions, enhancing open-set performance.

  1. GPT-Based Generative Caption Reconstruction Loss
  • Add a training objective that reconstructs masked caption tokens conditioned on visual features, using a pre-trained language model (e.g., GPT).

  • This forces the visual encoder to capture fine-grained semantic details and improves visual-textual reasoning, enabling the AI to better understand context and describe unseen objects.

  1. Target-Vocabulary-Assisted Training Protocol
  • Adopt a protocol where pseudo-labels are generated using a predefined target vocabulary (including novel classes) and used alongside base-class labeled data.

  • This allows the AI system to learn from both labeled and pseudo-labeled data, improving performance on both known and unknown categories without additional human effort.

What the Improved AI System Can Do:

  • Recognize and segment unseen object categories in images (e.g., rare animals, novel products) with significantly higher accuracy (e.g., +22 AP points on novel classes in OVIS) compared to prior methods.

  • Handle synonyms and out-of-vocabulary terms robustly—e.g., correctly segmenting a couch when trained only on sofa or settee.

  • Generate descriptive captions for images containing novel objects, improving tasks like image captioning, visual question answering, and human-AI interaction.

  • Operate in open-set panoptic segmentation with up to 18 PQ points improvement on unknown classes, enabling safer deployment in dynamic environments (e.g., autonomous driving with unexpected obstacles).

  • Maintain low computational overhead at inference (only 2.4% increase in GFLOPs), making it practical for real-time applications while still benefiting from rich multimodal training.

  • Reduce annotation costs by automatically generating pseudo-labels and captions, allowing AI systems to scale to new domains with minimal human intervention.

Sources

Related papers