Open Vocabulary Panoptic Segmentation With Retrieval Augmentation

arXiv:2601.12779 · cs.CV, cs.CL · Submitted 2026-01-19 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Open Vocabulary Panoptic Segmentation With Retrieval Augmentation".

Jane: The paper was written by Nafis Sadeq, Qingfeng Liu and Mostafa El-Khamy from East West University and Samsung Semiconductor, Inc..

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome back to the arXiv channel, everyone. Today we’re looking at a paper that’s been making waves in computer vision circles — it’s called “Open Vocabulary Panoptic Segmentation with Retrieval Augmentation.” Jane, I’ve got to say, even the title is a mouthful.

Jane: It really is, Tom, but let’s unpack that. Panoptic segmentation is the task where a computer looks at an image and labels every single pixel — not just “this is a car,” but “this is car number one, this is car number two.” It’s like giving the machine a complete map of everything in the scene.

Tom: Right, and the “open vocabulary” part is what makes it exciting. Traditional systems are trained on a fixed set of classes — say, eighty categories from COCO. But open vocabulary means you can ask the system to find anything, even things it never saw during training. Like, “find the llama” when the model was only trained on cats and dogs.

Jane: And that’s where the problem starts. These systems use something called CLIP, a vision-language model that’s really good at matching images to text. But when you crop out a masked region — a specific object in an image — and feed that to CLIP, the features don’t look like natural images anymore. There’s a domain shift.

Tom: A domain shift — that’s the gap between what the model learned on and what you’re actually giving it at test time. The authors, Nafis Sadeq, Qingfeng Liu, and Mostafa El-Khamy, they noticed this bottleneck and thought, “What if we don’t rely on CLIP alone? What if we retrieve similar examples from a database instead?”

Jane: So instead of asking CLIP to classify a masked region directly, you build a big database of masked features with known labels. At test time, you take your masked region, find the closest matches in that database, and use those labels to make your prediction.

Tom: Exactly. And that’s the “retrieval augmentation” part of the title. You’re augmenting the model’s knowledge by pulling in real examples. It’s like if I showed you a blurry photo and you couldn’t tell what it was, but I also showed you ten similar blurry photos with labels — suddenly you’d have a much better guess.

Jane: And the beauty is, you can keep adding to that database without retraining the model. New classes? Just add new image-text pairs. That’s huge for real-world deployment.

Tom: I want to bring in Lu from Tsinghua — Lu, you’ve been nodding along. What’s your take on the retrieval angle?

Lu: I think it’s a clever workaround for a fundamental limitation. CLIP is trained on whole images, not masked segments. So the authors are essentially saying, “Don’t fight the domain shift — embrace it.” Build a database of features that are also masked, so the query and the database are in the same feature space. That’s elegant.

Jane: And it works. They show on the ADE20k dataset, after fine-tuning on COCO, they jump from twenty-six point four PQ to thirty point nine PQ with their best backbone. That’s a four point five point absolute improvement.

Tom: PQ — panoptic quality — is the main metric here. It combines how well you label pixels and how well you separate instances. So a four point five point jump is substantial.

Lu: And what I find most exciting is that even in a fully training-free setup — no fine-tuning at all — they get a five point two point improvement. That suggests retrieval is a strong standalone signal, not just a crutch.

Jane: Lu, you’re right, and that’s what I want to dig into next — how they actually build that database and why it’s so effective. Stay with us.

Summary: Tom: We’re back with “Open Vocabulary Panoptic Segmentation with Retrieval Augmentation.” Jane, last segment we talked about the big idea — retrieving similar features from a database instead of relying on CLIP alone. But how do they actually build that database?

Jane: Great question. They use paired image-text data — basically, images with captions or labels. The pipeline has four steps. First, they run an open-vocabulary object detector called Grounding DINO to find bounding boxes for objects in the image. Then they feed those boxes to SAM, the Segment Anything Model, to get precise masks.

Tom: So SAM gives you the pixel-level mask, but you need the bounding box first to tell it what to focus on. Otherwise SAM might break a car into wheels, windows, and a body — which is useless for panoptic segmentation.

Jane: Exactly. Then they take the whole image and run it through CLIP to get dense features — a grid of feature vectors. And here’s the clever part: instead of running CLIP on each masked region separately, which would be computationally expensive, they just pool the dense features inside each mask. That gives them one feature vector per segment.

Lu: And that feature vector goes into the database along with the class label from the image-text pair. So you end up with a big lookup table of masked segment features and their associated classes.

Tom: And at inference time, they take a masked segment from the input image, use its feature as a query, and search the database for the nearest neighbors. The class labels of those neighbors vote on what the segment should be.

Jane: But they don’t stop there. They combine the retrieval score with the CLIP-based score. They have three classification paths: the in-vocabulary classifier for classes seen during training, the retrieval path, and the CLIP path. Then they blend them with hyper-parameters.

Meng: I’m Meng, by the way — I’ve been listening from the engineering side. My first thought is, how big is this database? And how fast is the retrieval? Because if you’re doing nearest neighbor search over millions of features at inference time, that could be a bottleneck.

Jane: That’s a fair concern. The paper doesn’t go deep into latency, but they do mention using approximate nearest neighbor search, which is standard for this kind of thing. And the database construction is offline — you build it once, then query it at test time.

Meng: And what about the quality of the database? If your detector misses objects or SAM produces bad masks, you’re putting garbage into the database.

Tom: That’s actually a point they address. They show that using Grounding DINO plus SAM gives much better masks than SAM alone. Without the bounding box prompts, SAM produces fragmented masks — a single object split into multiple pieces. With the detector guiding it, the masks are class-aware and much cleaner.

Lu: And there’s a robustness check too. They build the database using ADE20k training images and evaluate on ADE20k validation — that’s the same domain. But they also test with Google Open Images as the database, which is a completely different dataset. The improvement is smaller — about one point nine PQ instead of four point five — but it’s still positive. So the method doesn’t collapse when the database is out-of-domain.

Jane: That’s reassuring for real-world use, because you won’t always have a perfectly matched database. Meng, does that address your concern about practical deployment?

Meng: Partially. I’d still want to see inference time numbers, but the fact that it degrades gracefully is a good sign. And the training-free setup — where they don’t fine-tune anything — that’s really attractive for quick deployment.

Tom: And that’s exactly what we’re going to dig into next — the training-free results and how the mask proposal quality affects everything. Stick around.

Improvements: Tom: Welcome back. We’re still on “Open Vocabulary Panoptic Segmentation with Retrieval Augmentation.” Jane, last segment we covered how they build the database. Now let’s talk about what actually improves — and what doesn’t.

Jane: Right. The paper has two main setups. The first is cross-dataset fine-tuning — you train on COCO, evaluate on ADE20k. The second is fully training-free — no panoptic annotations at all. And in both cases, retrieval helps.

Tom: In the training-free setup, they compare against a CLIP-only baseline. With a ViT-large backbone, CLIP alone gets ten point nine PQ. Adding retrieval jumps that to sixteen point one PQ. That’s a five point two point absolute improvement.

Lu: And here’s the striking part — if you remove the CLIP classifier entirely and rely only on retrieval, you still get fifteen point eight PQ. That’s almost as good as the combined system. So retrieval alone is a stronger signal than CLIP alone for out-of-vocabulary classes.

Jane: That’s because of the domain shift we talked about. CLIP is trying to match a masked feature to a text embedding, and those two spaces are misaligned. But retrieval matches masked features to other masked features — same space, no shift.

Meng: So the retrieval path is essentially fixing the alignment problem by never leaving the visual domain. That makes sense. But what about the mask proposal quality? I remember you mentioned that SAM alone is bad.

Tom: Yeah, and they have a whole ablation on that. With ground truth masks — perfect masks — they get twenty-eight point four PQ. That’s a ceiling. With SAM alone, using grid-sampled point prompts, it crashes to seven point eight PQ. But with Grounding DINO providing bounding boxes to SAM, it recovers to sixteen point one PQ.

Jane: So the mask proposal is the bottleneck. Retrieval can only be as good as the masks you feed it. If the mask cuts an object in half or merges two objects, the feature you pool is wrong, and the retrieval will return mismatched neighbors.

Lu: That’s a classic chicken-and-egg problem in open vocabulary segmentation. You need good masks to classify, but you need to know the class to get good masks. The authors partially solve it by using an open-vocabulary detector to guide mask generation.

Meng: And in the cross-dataset setup, they fine-tune the mask proposal on COCO, which helps a lot. The Mask2former-based generator benefits from seeing annotated data, even if the classes are limited.

Tom: Right, and that’s where the four point five PQ improvement comes in — from twenty-six point four to thirty point nine with the ConvNeXt-large backbone. The fine-tuned mask generator gives cleaner proposals, and retrieval adds the out-of-vocabulary boost.

Jane: They also do hyper-parameter tuning for the ensemble weights — alpha, beta, gamma — and find the best balance is heavily weighted toward retrieval for unseen classes. That tells you the retrieval signal is trustworthy.

Lu: One thing I’d love to see explored is whether the database can be updated continuously. The paper mentions it’s training-free to add new classes — you just add new image-text pairs. But what about removing stale or noisy entries? That’s an open question.

Meng: And I’d want to know how the retrieval scales. If you go from twenty thousand segments to twenty million, does the accuracy hold? Approximate nearest neighbor search is fast, but the quality of the neighbors can degrade.

Jane: Those are great points, and they point to future work. But for now, the paper shows a solid, practical improvement. Let’s wrap up with our final thoughts.

Conclusion: Tom: We’ve reached the end of our discussion on “Open Vocabulary Panoptic Segmentation with Retrieval Augmentation.” Jane, give us the final takeaway.

Jane: The core idea is simple and powerful: when CLIP struggles with masked image features, don’t force it — retrieve similar masked features from a database instead. The authors show this consistently improves performance, whether you fine-tune or go fully training-free.

Tom: And the numbers back it up. On ADE20k, after COCO fine-tuning, they go from twenty-six point four to thirty point nine PQ. In the training-free setup, from ten point nine to sixteen point one PQ. Those are meaningful gains.

Lu: What I find most impactful is the modularity. You can swap in a better detector, a better mask generator, or a bigger database, and the retrieval module just gets better. It’s a plug-and-play improvement to any open vocabulary system.

Meng: And from an engineering standpoint, the fact that you can add new classes without retraining is a game-changer for deployment. You just update the database. That’s the kind of flexibility production systems need.

Jane: There are still open questions — mask proposal quality, database scale, and handling noisy entries. But the direction is solid. Retrieval augmentation is a practical, effective way to push open vocabulary segmentation forward.

Tom: And with that, we’ll say goodbye to “Open Vocabulary Panoptic Segmentation with Retrieval Augmentation.” Thanks to Nafis Sadeq, Qingfeng Liu, and Mostafa El-Khamy for their work. Next up, we’ve got a paper on diffusion models for video generation — that should be fun.

Jane: Until then, keep your eyes on the pixels, everyone. Tom and I will see you in the next episode.

Nafis Sadeq, Qingfeng Liu, Mostafa El-Khamy

East West University · Samsung Semiconductor, Inc.

cs.CV, cs.CL

Submitted: 2026-01-19

Updated: 2026-08-18

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 80/100

Key concepts

Panoptic Segmentation
This is a task where a computer labels every pixel in an image, assigning both a class (like 'car') and an instance ID (like 'car number one'). It creates a complete map of everything present in the scene.
Open Vocabulary
Traditional segmentation models are trained on fixed sets of classes. Open vocabulary allows the system to find and label objects it was never trained on, such as finding a specific llama when only cats and dogs were taught.
Retrieval Augmentation
Instead of relying solely on a vision-language model like CLIP for masked regions, this method builds a database of similar masked features. At test time, the system retrieves and uses labels from these similar examples to make predictions.

Terminology

Summary

Summary

This paper proposes RetCLIP, a retrieval-augmented method for open vocabulary panoptic segmentation that addresses the domain shift problem between natural image features and masked image features when using a CLIP vision encoder. The authors state: "In this work, we address the bottleneck mentioned above in the context of open vocabulary panoptic segmentation. In order to mitigate the domain shift between the natural image feature and the masked image feature, we propose RetCLIP, a retrieval-augmented approach for panoptic segmentation."

The method constructs a feature database from paired image-text data. The database construction process has four steps: object detection, mask generation, dense feature generation, and mask pooling. In object detection, an open vocabulary object detection method (Grounding DINO) is used to generate bounding boxes for each object instance. In mask generation, the input image and bounding box prompts are fed to Segment Anything Model (SAM) to generate class-aware masks. The authors note that "Even though SAM can generate masks without class-aware bounding boxes, the resulting masks often break up a single class (e.g. car) into multiple masks (e.g. wheel, car body, window). The class-aware masks generated in the previous step ensure that the SAM can generate high-quality masks for each class present in the image." In dense feature generation, CLIP extracts image-level dense features. In mask pooling, dense features associated with the whole image are used to generate mask-specific features, avoiding the need to encode each masked segment separately, which can be computationally expensive.

At inference time, the masked segment features from the input image are used as query keys to retrieve similar features and associated class labels from the database. The authors explain: Since both the retrieval key and retrieval target use a CLIP vision encoder on masked regions, the proposed approach does not suffer from the domain shift between the natural image feature and the masked image feature. The retrieval-based classification scores are combined with CLIP-based scores to produce the final output.

The system architecture is based on FC-CLIP with a frozen CNN-based CLIP backbone (ConvNeXt-Large variant from OpenCLIP). The mask proposal generator is based on Mask2former. The system has three classification paths: in-vocabulary classification (with fine-tuned linear projection on COCO), out-of-vocabulary classification via retrieval, and out-of-vocabulary classification via CLIP. The scores are combined using an ensemble with hyper-parameters α, β, and γ. The authors state: "Let's assume C is the set of classes for prediction and Ctrain is the set of classes in the fine-tuning dataset. Let siclip, siret, siiv be classification scores for class i using CLIP, retrieval and in-vocabulary classifier. The scores from the three classification pipelines are combined as follows, where α, β, γ are hyper-parameters." The formulas are: sioov = siret × γ + siclip × (1 − γ); si = sioov × α + siiv × (1 − α) if i ∈ Ctrain; si = sioov × β + siiv × (1 − β) if i ∉ Ctrain.

The authors evaluate in two settings: cross-dataset fine-tuning (fine-tuned on COCO, evaluated on ADE20k) and training-free (no panoptic segmentation annotations used). For the feature database, they use either the ADE20k train set or the Google Open Image dataset. They note: "In case any of the user-provided class names are missing in the feature database, we retrieve image samples for those input classes from a secondary fallback image dataset. We use ADE20k to construct our primary feature database and the Google Open Image dataset as a fallback. The label matching between datasets is performed with CLIP text embedding of class names with similarity score > 0.95."

Results in the cross-dataset setup show that with the ADE20k training set as a feature database, RetCLIP achieves an absolute improvement of +4.3 PQ and +4.5 PQ for CLIP-RN50x64 and CLIP-ConvNeXt-large backbone respectively over FC-CLIP. With the Google Open Image dataset as a feature database, the improvement is +0.9 PQ and +1.9 PQ for CLIP-RN50x64 and CLIP-ConvNeXt-large backbone respectively. The best result is 30.9 PQ, 19.3 mAP, and 44.0 mIoU on ADE20k with the CLIP-ConvNeXt-large backbone and ADE20k database.

In the training-free setup, RetCLIP achieves an absolute improvement of +3.7 PQ and +5.2 PQ for CLIP-ViT-base and CLIP-ViT-large backbone respectively over a CLIP-only baseline. The authors also find that retrieval-based classification alone outperforms CLIP-only baseline for out-of-vocabulary classes, with +2.7 PQ and +4.9 PQ for CLIP-ViT-base and CLIP-ViT-large backbone respectively, demonstrating that retrieval itself can be a strong baseline because it is robust to domain shift between natural image CLIP features and masked image CLIP features.

The authors also analyze the impact of mask proposal quality. With ground truth masks, the system achieves 28.4 PQ with CLIP-ViT-large backbone. Automatic mask generation with SAM alone performs poorly (7.8 PQ) because SAM is trained for interactive input with humans in the loop. Without human input, SAM masks are not class-aware. SAM may break up a single object into multiple fine masks. Using Grounding DINO to construct class-aware bounding boxes and feeding them to SAM improves PQ to 16.1.

Hyper-parameter tuning shows best performance with α = 0.4, β = 0.7, γ = 0.3.

The authors list their contributions as: "We proposed RetCLIP, a retrieval-augmented panoptic segmentation approach that tackles the domain shift between the natural image feature and masked image feature with respect to the CLIP vision encoder. The proposed approach can incorporate new classes in the panoptic segmentation system simply by updating the feature database in a fully training-free manner. The feature database can be constructed from paired image-text data which is widely available for thousands of classes. and We demonstrate that the proposed system can improve open vocabulary panoptic segmentation performance in both training-free setup (+5.2 PQ) and cross-dataset fine-tuning setup (+ 4.5 PQ, COCO→ADE20k)."

The authors conclude: "Even though the proposed method achieves reasonable performance in an open vocabulary setting, it remains vulnerable to the quality of mask proposal generation. Future work may focus on improving the quality of mask proposal generation for unknown classes."

Improvements for AI systems

Based on the paper, here are the specific improvements I can implement and the resulting capabilities of the improved AI system:

1. Retrieval-Augmented Classification Module

  • Build a feature database from paired image-text data using a pipeline of: Grounding DINO (object detection) → SAM (mask generation) → CLIP (dense feature extraction) → mask pooling

  • At inference, use masked segment features as query keys for approximate nearest neighbor search in the database

  • Combine retrieval-based scores with CLIP-based scores using an ensemble with tuned hyperparameters (α=0.4, β=0.7, γ=0.3)

2. Domain-Shift Mitigation for Masked Features

  • Replace the naive approach of encoding masked regions with CLIP (which suffers from domain shift between natural images and masked images) with retrieval from a database of masked segment features

  • Since both query and database features come from masked regions, the domain shift is eliminated

3. Training-Free Class Expansion

  • Enable adding new classes without any retraining by simply updating the feature database

  • Use a fallback image dataset (e.g., Google Open Images) when a class is missing from the primary database, with label matching via CLIP text embedding similarity > 0.95

4. Improved Mask Proposal Generation

  • For training-free setup, replace naive SAM mask generation (which breaks objects into fine parts) with Grounding DINO bounding boxes → SAM pipeline for class-aware masks

Quantitative Performance Gains (on ADE20k, fine-tuned on COCO):

  • +4.5 PQ (26.4 → 30.9) with CLIP-ConvNeXt-large backbone

  • +2.5 mAP (16.8 → 19.3)

  • +10.0 mIoU (34.0 → 44.0)

Training-Free Performance Gains (on ADE20k):

  • +5.2 PQ (10.9 → 16.1) with CLIP-ViT-large backbone

  • +3.4 mAP (6.9 → 10.3)

  • +8.4 mIoU (13.8 → 22.2)

Robustness to Database Domain Shift:

  • Even when using Google Open Images (different domain from ADE20k test data), still achieves +1.9 PQ improvement over baseline

Key Operational Capabilities:

  • Can segment arbitrary user-specified classes without retraining

  • Handles classes never seen during training by retrieving similar visual features from the database

  • Works with both frozen CLIP backbones (RN50x64, ViT-base, ViT-large, ConvNeXt-large)

  • Degrades gracefully: retrieval alone (without CLIP) still outperforms CLIP-only baseline by +2.7 to +4.9 PQ, proving robustness

Implementation Notes:

  • All components (Grounding DINO, SAM, CLIP) are frozen; only the in-vocabulary linear projection is fine-tuned

  • The feature database construction is fully automated from image-text pairs, requiring no pixel-level annotations

  • The system can be extended to new datasets by rebuilding the database from available paired data

Related papers