Open-World Panoptic Segmentation
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "Open-World Panoptic Segmentation".
Tom: Perception is a key building block of autonomously acting vision systems such as autonomous vehicles, and this paper tackles open-world panoptic segmentation by proposing Con2MAV,
Jane: First, who's behind it and why it matters.
Paper summary: Tom: So we're talking about this paper, "Open-World Panoptic Segmentation," and what they’re trying to do there is tackle the problem of seeing things that you’ve never seen before when you're driving or walking in the real world.
Jane: Exactly. The main idea here is discovering new semantic categories and new object instances at test time while keeping everything consistent among those categories as they are discovered incrementally, which they call Con2MAV one <ref:2412.12740#pg0,discovering new semantic categories and new object instances at test time while>.
Lu: It’s really interesting because they are extending a previous work, ContMAV, which was focused on open-world semantic segmentation to now also predict individual instances one <ref:2412.12740#pg0,previous work, ContMAV, which was>.
Meng: So it's not just about knowing what the stuff is, like "this is a chair," but also figuring out that "this specific chair" is an instance of that category. That sounds complex for a real system to handle in real-time.
Lalam: From my perspective, if the AI can identify a brand new type of object it has never seen before, it means our culture and how we categorize things could adapt much faster than we currently expect one <ref:2412.12740#pg0>.
Tom: It matters because autonomous systems need this ability to operate safely in unconstrained scenarios where they can't just rely on what they were trained on previously.
Jane: They claim that Con2MAV achieves state-of-the-art results across several datasets, including SegmentMeIfYouCan, COCO, BDDAnomaly from CAOS, and PANIC one <ref:2412.12740#pg0,achieves state-of-the-art results>.
Lu: That’s a lot of work they put into testing it on different kinds of open-world tasks to show it works well in various domains.
Meng: I'm curious about the architecture since that seems like the core mechanism for how it handles all this stuff. What kind of neural network are they using?
Tom: They propose Con2MAV, which is a fully convolutional neural network with one encoder and three decoders one <ref:2412.12740#pg0>. The encoder uses a ResNet34 but they've changed the standard blocks to a NonBottleneck-1D block one <ref:2412.12740#pg0>.
Paper summary: Jane: That change in the encoder, replacing standard convolutions with those three times one and one times three convolutions, is meant to incorporate contextual information more effectively through a pyramid pooling module at the end of that part one.
Lu: Then you've got three decoders, all structured similarly using SwiftNet modules and NonBottleneck-1D blocks, with final upsampling using nearest-neighbor and depth-wise convolutions which they say significantly reduces the computation needed one <ref:2412.12740#pg0>.
Meng: So it’s a lot of layers working together to figure out what’s there. What does the first decoder actually do?
Jane: The first one, the semantic decoder, targets semantic segmentation and also tries to build a unique class descriptor for each known category one <ref:2412.12740#pg0>.
Tom: They use a weighted cross-entropy loss for that closed-world segmentation, and they add another loss function called Lfeat at the pre-logit level to build those class prototypes one <ref:2412.12740#pg0>.
Lu: Then there's the second decoder, the contrastive decoder, which focuses on anomaly segmentation by mapping known class features onto a unit hypersphere while pushing unknown features to zero one <ref:2412.12740#pg0>.
Jane: That contrastive loss they use is designed to separate what’s familiar from what’s new in the feature space using Lcont one <ref:2412.12740#pg0>.
Meng: And the third decoder, the instance decoder, handles class-agnostic instance segmentation using vector fields-inspired loss functions like Lovasz Hinge loss one <ref:2412.12740#pg0>.
Tom: The authors point out that they use pre-logit features instead of pre-softmax features for computing those losses, and they found that this boosts performance by about four to five percent on both metrics one <ref:2412.12740#pg0>.
Jane: They also mention that the learned thresholds bring an even further improvement when paired with the pre-logit approach, probably because those class descriptors are more robust one <ref:2412.12740#pg0>.
Lu: For the final panoptic segmentation, they use HDBScan to cluster predictions in the "thing" areas, which includes all known classes and all unknown categories one <ref:2412.12740#pg0>.
Meng: So for someone who just uses an AI to look around, what does this mean practically for their daily experience? Does it mean fewer errors or something bigger?
Tom: It means that when you use a vision system in a completely new environment, it can actually discover and label the objects you haven't seen before instead of just failing or guessing one <ref:2412.12740#pg0>.
Paper summary: Jane: The paper introduces the PANIC dataset as an open-world panoptic segmentation test set, which has eight hundred images and includes fifty-eight unknown categories one.
Lu: Having those fifty-eight unknown classes in the test set really tests how robust the system is when it has to invent new concepts on its own one.
Meng: The paper also includes benchmarks for anomaly segmentation, where they measure metrics like AUPR and FPR95 for that part of the task one <ref:2412.12740#pg0>.
Tom: Overall, this work shows state-of-the-art results on various open-world segmentation tasks while still performing competitively on the known classes one <ref:2412.12740#pg0,open-world segmentation tasks while still performing competitively on the known>.
Jane: The authors also noted a limitation, which is that the model relies on several manually tuned thresholds for predicting unknown things, and that's something they acknowledge one <ref:2412.12740#pg0>.
Lu: That reliance on manual tuning is a real constraint when you want a system that works completely autonomously without human intervention in setup one <ref:2412.12740#pg0>.
Meng: So the paper tackles the hard part of discovering new categories at test time while trying to keep the known ones stable, but we have to remember that tuning those thresholds still needs human input one <ref:2412.12740#pg0>.
Tom: That's where it points us next: how do we make those threshold determinations more automatic or robust across different environments?
Jane: This leads us into the conclusion of this discussion about "Open-World Panoptic Segmentation," where the authors discuss the title and their implications for future research one <ref:2412.12740#pg0>.
Lu: The implication is that we can build vision systems that are truly adaptable to novel situations in the real world, not just those they were explicitly trained on one <ref:2412.12740#pg0>.
Meng: For practical AI development, this suggests a path toward deploying systems in unpredictable environments where new objects pop up unexpectedly without needing a full retraining cycle one <ref:2412.12740#pg0>.
Tom: So this paper is about moving vision from closed-world to open-world tasks by letting the system discover and maintain consistency among what it finds on the fly one <ref:2412.12740#pg0>.
Conclusion: Tom: So we've seen how Con2MAV handles discovering new things in open-world vision, and now we're at the end of this paper discussing what it actually means for us.
Jane: They call it "Open-World Panoptic Segmentation," which basically means the system can handle both recognizing what’s around us and knowing exactly where every single object is, even if they are completely new to it.
Lu: It’s about taking a vision system that only knows things from its training data and making it capable of figuring out totally novel classes on the fly while keeping track of all the instances.
Meng: So for someone just listening to this, what does that actually translate to when they're using a self-driving car or even just looking at a live camera feed?
Lalam: It means that instead of getting stuck and saying "I don't know what that is," the AI can start naming and tracking brand new things it sees in real-time.
Tom: Exactly. The authors are really pushing the boundaries here by proposing this method for discovering categories at test time while trying to keep everything consistent among those new discoveries.
Jane: It’s a big step because it moves vision systems away from just recognizing familiar stuff and toward actually understanding an entire world that keeps changing.
Lu: They've also introduced a new test set called PANIC, which has fifty-eight unknown categories, so they are really testing the system's ability to invent concepts on its own in a very diverse way.
Meng: I see it as making those autonomous systems much more robust when they encounter unexpected situations out there in the real world.
Lalam: It means our AI could adapt much faster to new objects or environments without needing a massive retraining effort every single time something different pops up.
Tom: That adaptability is what’s really exciting here, moving us closer to having vision systems that can truly operate in any unconstrained setting they might encounter.
Jane: But as the authors pointed out, they do rely on certain thresholds for deciding if something is new or known, which is a practical thing to keep in mind.
Lu: Yeah, figuring out how to make those threshold decisions really automatic and reliable across different kinds of real-world conditions is the next big challenge they're facing.
Tom: That's the big question we need to follow up on next. How do we get that system to figure out what’s new without needing a human hand just to set the rules?
Matteo Sodano, Federico Magistri, Jens Behley, Cyrill Stachniss
University of Bonn
cs.CV, cs.RO
Submitted: 2024-12-17
Updated: 2026-10-05
Comments: Accepted at IJRR
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 85/100
The gist: Perception is a key building block of autonomously acting vision systems such as autonomous vehicles, and this paper tackles open-world panoptic segmentation by proposing Con2MAV, an approach that
Key concepts
- Open-World Panoptic Segmentation
- This task requires a system to segment every pixel in an image into two parts: semantic labels (what the object is) and instance IDs (which specific object it is). Open-world means the system must correctly identify objects and categories it has never seen during training, enabling adaptation to novel situations.
- Semantic Decoder
- This decoder performs standard semantic segmentation to classify pixels into known categories. Crucially, it also builds a unique class descriptor or prototype for each known category by applying a specific loss function at the pre-logit level, helping the system recognize new classes.
- Contrastive Decoder
- This component focuses on anomaly detection. It maps features of pixels belonging to known classes onto a unit hypersphere while mapping unknown class features to zero. This helps distinguish between familiar objects and novel anomalies in the scene.
- Instance Decoder
- This decoder handles instance segmentation, aiming for class-agnostic identification of individual objects. It uses vector fields-inspired loss functions to guide offset predictions toward object centroids, ensuring accurate boundaries for all detected instances.
Terminology
Summary
Perception is a key building block of autonomously acting vision systems such as autonomous vehicles, and this paper tackles open-world panoptic segmentation by proposing Con2MAV, an approach that discovers new semantic categories and object instances at test time while enforcing consistency among discovered categories. This work matters because it addresses the crucial need for autonomous systems to operate safely in unconstrained real-world scenarios where novel situations and objects must be identified without prior training knowledge.
The gist: Con2MAV is a fully convolutional neural network that achieves state-of-the-art performance on open-world segmentation tasks, extending ContMAV in order to predict also instances and not only semantic categories, while still providing compelling closed-world performance in Sec. 4<ref:2412.12740#pg6>.
How it works
Con2MAV is a novel method for open-world panoptic segmentation that jointly addresses anomaly segmentation, open-world semantic and open-world panoptic segmentaton, and achieves state-of-the art results on several public datasets, such as SegmentMeIfYouCan [10], COCO [44], BDDAnomaly from CAOS [30], and SUIM [35]<ref:2412.12740#pg3>. It is an extension of the previous work ContMAV, which was developed for open-world semantic segmentation<ref:2412.12740#pg6>.
The architecture of Con2MAV consists of one encoder and three decoders<ref:2412.12740#pg6>. The encoder uses a ResNet34 [29] encoder, where the standard ResNet block is replaced with the NonBottleneck-1D block [68], which replaces all 3 × 3 convolution by a sequence of 3 × 1 and 1 × 3 convolutions with a ReLu in between<ref:2412.12740#pg6>. The architecture also incorporates contextual information by means of a pyramid pooling module [91] at the end of the encoding part<ref:2412.12740#pg6>.
The three decoders have the same structure and are composed of three SwiftNet modules [61] with NonBottleneck-1D blocks, and two final upsampling modules based on nearest-neighbor and depth-wise convolutions [14], that have the major advantage of substantially reducing the computations needed<ref:2412.12740#pg6>. The first decoder is called “semantic decoder”, which targets semantic segmentation and additionally tries to build a unique class descriptor for each known category<ref:2412.12740#pg6>.
How it works (continued)
The semantic decoder aims to perform semantic segmentation and it leverages the standard weighted cross-entropy loss for closed-world semantic segmentation: Lsem = − 1 / omega X p∈omega ωk t ⊤ p log σ(f p) <ref:2412.12740#pg7>. Besides the standard closed-world semantic segmentation, the goal with the semantic decoder is also to build a unique class descriptor, or class prototype, for each known class<ref:2412.12740#pg7>. This is achieved by applying a loss function Lfeat at the pre-logit level given by Lfeat = 1 / omega X K k=1 X p∈omegak [lp − µe−1k] squared / σe−1k squared <ref:2412.12740#pg7>.
The second decoder, called “contrastive decoder”, targets anomaly segmentation and tries to map the features of pixels belonging to known classes on a unit hypersphere, while mapping features of pixels belonging to unknown classes to 0<ref:2412.12740#pg6>. It uses the contrastive loss Lcont = − X K k=1 log exp (lk ⊤ µ¯e−1k / τ) PK i=1 exp (lk ⊤ µ¯e−1i / τ) <ref:2412.12740#pg7>.
The third decoder, called “instance decoder”, targets class-agnostic instance segmentation by means of vector fields-inspired loss functions [83]<ref:2412.12740#pg6>. It uses the Lovasz Hinge loss Loff = 1 / C X C j=1 Lov´asz(FCj, GCj) <ref:2412.12740#pg6>.
How it works (continued)
The post-processing for anomaly segmentation fuses the outputs of the semantic and contrastive decoders<ref:2412.12740#pg6>. The semantic decoder yields closed-world semantic segmentation, but simultaneously builds individual class descriptors in the pre-logit space for each known class<ref:2412.12740#pg7>. A simple 1σ bound is used as a decision mechanism to determine whether a pixel is unknown or not<ref:2412.12740#pg7>.
The post-processing for open-world semantic segmentation involves constructing an open-world confusion matrix where the number of rows K˜ refers to the newly-discovered classes, and the number of columns K is fixed and corresponds to the ground truth classes<ref:2412.12740#pg7>.
The post-processing for open-world panoptic segmentation uses HDBScan [51] to cluster the offset predictions, executing this clustering only in the “thing” areas, which means all the known classes that are known to have instances, and all the unknown categories<ref:2412.12740#pg7>.
How it works (continued)
The approach relies on pre-logit features instead of pre-softmax features for computing the loss functions of the semantic and contrastive decoder, which is shown to boost performance by 4 − 5% on both metrics<ref:2412.12740#pg14>. The learned thresholds bring a further improvement when paired with the pre-logit, probably because the class descriptors are more robust and the post-processing manages to better separate the newly-discovered classes<ref:2412.12740#pg14>.
The instance decoder is optimized with a weighted sum of loss functions Lins = w5 Loff + w6 Ldiv + w7 Laux div + w8 Lcurl + w9 Laux curl <ref:2412.12740#pg8>. The divergence and curl loss functions are inspired by the concept of vector field in physics, aiming to make the offset vectors point toward their corresponding centroid and to ensure no rotational behavior<ref:2412.12740#pg8>.
The final post-processing for open-world panoptic segmentation uses HDBScan [51] to cluster the offset predictions, and this result is filtered by the semantic prediction in order to enforce consistency<ref:2412.12740#pg8>.
How it works (continued)
The paper proposes a new dataset called PANIC, which is an open-world panoptic segmentation test set in the autonomous driving context, composed of 800 images and containing 58 unknown categories<ref:2412.12740#pg9>. It provides pixel-wise annotations of semantic classes and individual object instances<ref:2412.12740#pg4>.
How it works (continued)
The proposed benchmarks include four benchmarks accompanied by public Codabench competitions [87], one for each of the segmentation tasks described in Sec. 3<ref:2412.12740#pg6>. For anomaly segmentation, pixel-level metrics include the area under the precisionrecall curve (AUPR) and false positive rate at 95% true positive rate (FPR95)<ref:2412.12740#pg9>.
How it works (continued)
The experimental evaluation shows that Con2MAV achieves state-of-the-art results on several datasets on all kinds of open-world segmentation tasks, while performing competitively on the known classes<ref:2412.12740#pg7>. The model is shown to achieve compelling results on all three datasets, consistently outperforming the baselines<ref:2412.12740#pg7>.
REFERENCES
[1] J. Ackermann, C. Sakaridis, and F. Yu, “Maskomaly: Zero-shot mask anomaly segmentation,” in Proc. of British Machine Vision Conference (BMVC), 2023<ref:2412.12740#pg16>.
[6] H. Blum, P. E. Sarlin, J. Nieto, R. Siegwart, and C. Cadena, “Fishyscapes: A benchmark for safe semantic segmentation in autonomous driving,” in Proc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition Workshops, 2019<ref:2412.12740#pg9>.
[7] H. Blum, P. E. Sarlin, J.
Improvements for AI systems
- Bold header: Improved Open-World Panoptic Segmentation Capability
This system can discover new semantic categories and new object instances at test time, while enforcing consistency among the categories that we incrementally discover,
directly addressing the core problem of open-world panoptic segmentation by combining semantic and instance discovery.
- Bold header: Robust Novel Class Discovery
The model can discover not only novel and consistent semantic classes, but also individual objects within these classes,
moving beyond prior work by achieving a holistic understanding of the scene, since every pixel belongs to a semantic class and will have an instance ID (when applicable), irrespectively of it being known or unknown.
- Bold header: Improved Closed-World Performance Under Open-World Conditions
The system maintains high reliability on familiar data by leveraging learned descriptors, as the paper shows that the performance of our approach on the closed-world categories is not harmed by the open-world nature of the CNN,
ensuring compelling closed-world performance.
- Bold header: Enhanced Anomaly Segmentation Accuracy
The system can perform state-of-the-art anomaly segmentation, achieving high AUPR and PPV on challenging datasets like PANIC, which is crucial for safely identifying novel or unexpected objects in real-world scenarios.
- Bold header: Unified Nomenclature for Open-World Tasks
The proposed non-ambiguous nomenclature
aims to unify the nomenclature for open-world tasks, hoping that this would make open-world segmentation more common and accessible in the future,
improving communication and standardization across the field.
Sources
- Outlier detection by ensembling uncertainty with negative objectness
- Learning Confidence for Out-of-Distribution Detection in Neural Networks
- OoDIS: Anomaly Instance Segmentation and Detection Benchmark
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models