Open-World Hierarchical Perception: Taxonomic Abstraction over Class-Agnostic Proposals for the Safe Handling of Out-of-Vocabulary Road Objects

arXiv:2608.07577 · cs.CV, cs.AI, cs.RO · Submitted 2026-08-04 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Open-World Hierarchical Perception: Taxonomic Abstraction over Class-Agnostic Proposals for the Safe Handling of Out-of-Vocabulary Road Objects".

Jane: The paper was written by Felix Schaller from.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome back to the arXiv radio hour, folks. I'm Tom, and alongside me is the brilliant Jane, and we're diving into a paper that's got a mouthful of a title: "Open-World Hierarchical Perception: Taxonomic Abstraction over Class-Agnostic Proposals for the Safe Handling of Out-of-Vocabulary Road Objects."

Jane: Tom, that title is a whole research roadmap in itself, but let's break it down for our listeners. The core idea is about self-driving cars and what happens when they see something they've never been trained on. You know, a horse-drawn carriage, a piece of tire on the highway, a deer crossing at dusk.

Tom: Right, and normally, a car's brain—the object detector—has a fixed list of things it knows: cars, trucks, pedestrians, bicycles. If it sees a deer, it doesn't have a "deer" box to check. So what does it do?

Jane: Exactly, and that's the dangerous part. It either forces the deer into the closest box it knows, like "pedestrian," or it just ignores it completely. Both are terrible outcomes for safety. This paper proposes a smarter way: instead of forcing a specific label, the car can say, "I see a living being," or even just, "I see an unknown obstacle right here."

Tom: So it's about giving the car permission to be vague but correct, rather than specific and wrong. I love that framing. And the paper is part of a series, right? The authors have been building on this hierarchical taxonomy idea for a while.

Jane: That's right. This is the third paper in the series. The earlier ones proved the concept on objects that a normal detector already finds. But this new one tackles the real open world—objects the detector would never even box up in the first place. That's a huge leap.

Tom: And that's what we're going to dig into today. We've got our senior researcher Lu and our engineer Meng joining us later to really poke at the methodology and the practical implications.

Jane: For now, I want listeners to hold onto one number: the paper says a flat, closed-set detector is confidently wrong one hundred percent of the time on out-of-vocabulary objects. We're talking about a system that is always, always wrong when it meets the unknown. That's the problem this paper is trying to solve.

Tom: And it does it with a pretty clever trick: a safety floor in the taxonomy. We'll get into that. But first, let's just appreciate the ambition of the title. It's not just about detecting; it's about perceiving safely.

Jane: Exactly. And stick around, because in the next segment, we're going to look at the actual summary and the key results that make this more than just a nice idea.

Tom: You're listening to the arXiv hour. Don't go anywhere.

Summary: Tom: We're back, and we're still on "Open-World Hierarchical Perception." Jane, you set the stage with the problem—the one hundred percent confident-wrong rate. Now let's talk about what the authors actually did to fix it.

Jane: So the fix is a three-stage pipeline: propose, classify, validate. First, they use a class-agnostic segmenter—think of it as a tool that finds blobs in an image without knowing what they are. That gives them regions for everything, including the deer.

Tom: And then they classify those regions using a hierarchical taxonomy. Instead of a flat list, they have a tree: "Vehicle" branches down to "Car" and "Truck," "Living Being" branches down to "Horse" and "Pedestrian."

Jane: Right, and the clever part is the abstraction rule. They don't just pick the most likely leaf on the tree. They aggregate the evidence upward. If the system is torn between "Horse" and "Cow," it might just say "Living Being" instead of guessing wrong.

Tom: And if the evidence is too weak even for that? It says "UNKNOWN OBSTACLE." That's the safety net. The paper calls it a safety floor—the coarsest level that's still useful for planning. Below that, you don't get a label; you get a localized warning.

Jane: And the results are striking. They held out seven classes from the taxonomy—like trucks and horses—and tested on two hundred thirty-five real objects. The flat head was confidently wrong on all of them, and thirty-seven percent of those errors were in the wrong super-category. That means it called an animal a vehicle.

Tom: Meanwhile, the hierarchical layer—they call it HOWC—emitted zero confident wrong specific labels. Zero. And it safely handled ninety-four percent of the objects, either with the right super-category or an honest "unknown."

Jane: But here's the honest part they emphasize: the correct super-category was only recovered twenty-six percent of the time. The other sixty-nine percent were flagged unknown. So this is a safety win, not a recognition win. They're trading specificity for the elimination of categorical mistakes.

Tom: That's a trade I'd take on the road. A planner can brake for "unknown obstacle ahead," but it can't plan for a "sedan" that's actually a standing horse.

Jane: Exactly. And the paper is very clear that this is a safety result, not a specificity one. They're not claiming to recognize the deer better; they're claiming to not lie about it.

Tom: So the summary is: propose everything, classify carefully, and when in doubt, say "I don't know" out loud. That's a huge philosophical shift for autonomous driving.

Jane: It is. And it sets up the next question: how do they actually make this work in practice? That's where the feasibility study comes in, and that's our next segment.

Tom: Stay with us.

Improvements: Tom: Welcome back. We've covered the problem and the core solution. Now, Jane, the paper doesn't just stop at the taxonomy—it also explores how to feed this system in the open world. What improvements are they suggesting?

Jane: Right, Tom. The big improvement is the front-end. A normal detector won't box a deer, so they need a different way to get regions. They use class-agnostic segmentation—MobileSAM, specifically—to propose everything in the scene, regardless of category.

Tom: And that gives them recall, but it also gives them a ton of garbage. The paper found that on anomaly imagery, the closed detector found nineteen detections, but MobileSAM added one hundred forty-eight more regions. And a lot of those were background—vegetation, sky, road.

Jane: Exactly. So they needed a precision filter. They tested three signals: segmentation, appearance-based out-of-distribution scoring, and monocular depth. And this is where the paper gets really honest, because two of those signals partially failed.

Tom: The appearance-based OOD score was a negative result. They tried to use the margin between "this is a known thing" and "this is background" to filter out junk, and it just didn't work. The margins overlapped almost completely.

Jane: That's a great finding, actually. It tells us that a 2D crop alone can't tell you if something is real or a textureless background. They even had a case where an aircraft matched "plain background" more than "aircraft." The signal just isn't there.

Tom: And monocular depth? That worked better, but with limits. It cleanly separated foreground from background, which is a great precision cue. But it's scale-limited—small or distant objects are just "not assessable," and a compact animal can read as flat.

Jane: So the improvement they're suggesting is a combination: segmentation for recall, geometry for precision, and the taxonomy for semantics. No single 2D cue is enough. That's the synthesis.

Tom: And that's a really practical takeaway. It means the architecture needs to be modular—you can swap in better proposers and better validators without touching the semantic core.

Jane: Right. And the paper frames this as a feasibility study, not a finished product. They're saying, "Here's what works, here's what doesn't, and here's how they fit together." That's the kind of honest engineering we need more of.

Tom: So the improvements are about building a dependable open-world front-end, not just a smarter classifier. And that leads us to the actual evaluation—the leave-classes-out benchmark—which is the real meat of the paper.

Jane: And that's our next segment. We'll get into the numbers that make this paper stand out.

Tom: Don't move.

First Page: Tom: We're back on the arXiv hour, still deep in "Open-World Hierarchical Perception." Jane, let's zoom out and look at the first page of this paper, because it sets the tone for everything.

Jane: Absolutely, Tom. The abstract is a masterclass in honest claims. It starts by stating the problem: a closed-set detector on an out-of-vocabulary object can only force a wrong label or drop the object. Then it states the contribution: a layer that never makes a confident categorical mistake on an out-of-vocabulary object.

Tom: And they're explicit about the cost. They say, "We are explicit that this is a safety result, not a specificity one." That's rare in academic writing. They're not overselling.

Jane: It's refreshing. They report the correct super-category is recovered only twenty-six percent of the time, and sixty-nine percent are conservatively flagged unknown. They're putting the limitation right in the abstract, not burying it in the discussion.

Tom: And the framing is all about functional safety. The wrong label isn't just an accuracy issue—it's a categorical error that attaches the wrong size, mass, and behavior model to a real obstacle. That's a planning nightmare.

Jane: Right. If the car thinks a standing horse is a sedan, it plans for a sedan's braking distance and mass. That's how you get a catastrophic failure. The paper's whole point is that "I don't know" is a valid, actionable output.

Tom: And they mention the series context—this is the third paper. The earlier ones established the taxonomy on closed-detector boxes. This one takes it open-world. So it's a natural progression, not a one-off.

Jane: And the first page also introduces the key mechanism: the safety floor. Each branch of the taxonomy has a floor—the coarsest level that's still actionable. Below that, you get UNKNOWN OBSTACLE, not a useless generic "object" label.

Tom: That's a subtle but important design choice. A generic "object" label is useless for planning. But "unknown obstacle at this location" is actionable—the car can slow down, stop, or swerve.

Jane: Exactly. And that's the kind of thinking that makes this paper stand out. It's not just about classification accuracy; it's about what the output means for the downstream system.

Tom: So the first page sets up a safety-first philosophy, an honest evaluation, and a concrete mechanism. That's a strong opening. And it leads us to the conclusion, where we'll wrap up the big picture.

Jane: Let's get there.

Conclusion: Tom: And we're at the close of our discussion on "Open-World Hierarchical Perception: Taxonomic Abstraction over Class-Agnostic Proposals for the Safe Handling of Out-of-Vocabulary Road Objects." Jane, let's pull it all together.

Jane: Tom, the big picture is this: we've got a system that converts one hundred percent confident-wrong into zero percent confident-wrong on out-of-vocabulary objects. That's the headline. And it does it by allowing the system to say "living being" or "unknown obstacle" instead of forcing a specific guess.

Tom: And the cost is a sixty-nine percent abstention rate. That's a lot of "I don't knows." But the paper is clear that in a safety context, that's the right trade. A flagged unknown is recoverable; a confident wrong label is not.

Jane: Right. And they're already pointing to the future: metric geometry from LiDAR or stereo, motion cues, and a hierarchical data engine that turns the layer's own open-world output into training data for a taxonomy-native detector.

Tom: That last part is exciting. The system generates a corpus with fine labels where it's confident, coarse labels where it's not, and unknowns as a review queue. That's a self-improving loop.

Jane: And we should mention the limitations they're honest about: the evaluation isolates classification given ground-truth boxes, and the proposal front-end isn't yet dependable in pure 2D. Monocular depth is relative, not metric.

Tom: So it's not a finished product, but it's a solid foundation. And the philosophical shift is huge: perception systems should be allowed to be less specific but still correct.

Jane: That's the takeaway for our listeners. This paper isn't just about better detection; it's about a different relationship with uncertainty. And that's a change that could make autonomous driving genuinely safer.

Tom: Well said, Jane. We've covered the title, the summary, the improvements, and the first page. We're ready to say goodbye to this one and look ahead to the next paper on the arXiv.

Jane: Thanks for listening, everyone. We'll be back with more research that's shaping the future.

Tom: Until next time, keep your eyes on the road—and your taxonomies hierarchical.

Felix Schaller

cs.CV, cs.AI, cs.RO

Submitted: 2026-08-04

Updated: 2026-08-11

Comments: 6 pages, 4 figures, 1 table. Third paper in a series; v1 archived at Zenodo, doi:10.5281/zenodo.21593472. Code: https://github.com/freshNfunky/IE2025-Research-Paper

Code: https://github.com/freshNfunky/IE2025-Research-Paper

License: http://creativecommons.org/licenses/by-nc-sa/4.0/

Importance score: 46/100

The gist: The paper presents an open-world perception layer for autonomous driving that replaces a flat, closed-set object detector label space with a hierarchical taxonomy and a runtime abstraction rule,

Key concepts

Flat, Closed-Set Detector
A normal object detector has a fixed list of known objects (like cars or trucks). When it encounters something it hasn't been trained on, this system forces that unknown item into one of the existing boxes, which can lead to dangerous misclassification.
Hierarchical Taxonomy
This is a tree structure for classifying objects. Instead of a flat list, categories branch out (e.g., 'Living Being' branches into 'Horse' and 'Pedestrian'). This structure allows the system to aggregate evidence upward instead of making a single, potentially wrong guess.
Safety Floor
This is the coarsest level in the taxonomy that remains useful for planning. If evidence is too weak even for a specific label, the system outputs 'UNKNOWN OBSTACLE.' This provides an actionable safety net rather than a useless generic object label.
Class-Agnostic Segmenter
This tool finds regions (blobs) in an image without knowing what they are. It proposes regions for everything, including novel objects, which is the first step before applying the hierarchical classification layer.

Terminology

Summary

The paper presents an open-world perception layer for autonomous driving that replaces a flat, closed-set object detector label space with a hierarchical taxonomy and a runtime abstraction rule, extended to operate on class-agnostic region proposals rather than only on boxes produced by a closed detector. The authors state: "This paper takes the layer open-world. We (i) place the taxonomic abstraction layer on top of class-agnostic region proposals so that objects the closed detector never boxes can still be classified or flagged; (ii) report a feasibility study of three open-world signals (class-agnostic segmentation, appearance-based out-of-distribution scoring, and monocular depth) that shows why no single two-dimensional cue is sufficient and how they compose; and (iii) run the evaluation the earlier papers could not: a ground-truth leave-classes-out benchmark on real annotated objects."

The method is a three-stage pipeline: propose, classify, validate. Objects are organized in a directed tree whose leaves are concrete classes (Sedan, Cyclist, Horse) and whose internal nodes are safety-relevant abstractions (Vehicle, Living Being, Static Object). Each branch declares a floor: the coarsest level that is still actionable for planning. Abstraction is allowed down to the floor; anything below it is reported as UNKNOWN OBSTACLE rather than as a too-generic 'object'. The closed front-end runs a pretrained YOLO detector at high recall, and the open-world front-end adds MobileSAM for class-agnostic segmentation, so an untrained object still yields a region to classify. For each region, CLIP image embeddings are compared to text prompts for every taxonomy leaf, giving a softmax distribution over leaves. Rather than taking the arg-max leaf, the method aggregates leaf mass upward: "the score of an internal node is the sum of the mass of its descendant leaves. Starting at the root, we descend to a child only while that child concentrates at least a commit fraction m of the local mass; where the mass splits, we stop and report the current node, a justified abstraction. If even the best leaf's absolute similarity is below a floor smin, or descent halts below the branch safety floor, the region is reported as UNKNOWN OBSTACLE." A monocular depth map serves as an optional precision filter to suppress background regions.

The feasibility study reports three findings. First, class-agnostic proposals deliver recall but poor precision: "the closed YOLO front-end produced 19 detections on a sample where MobileSAM added 148 further regions the detector had missed. The hierarchical UNKNOWN gate correctly filtered 112 of these as unknown... but the 36 that received a category were mostly background (vegetation labelled 'Living Being'). Second, appearance-based OOD scoring is a negative result: the margins of implausible, plausible-out-of-taxonomy and good in-taxonomy cases overlap almost entirely, and the best threshold that catches all implausible cases also breaks a third of the good ones. Tellingly, a full-frame aircraft matched 'a plain textureless background' more strongly than 'an aircraft'. Third, monocular depth provides real but scale-limited signal: a per-region flatness/foreground test on a monocular depth map cleanly separated foreground from background... and, behind an assessability gate, read large upright objects as three-dimensional while holding small or distant regions as 'not assessable' rather than misflagging them. But relief is shape-sensitive (a compact animal read as flat), and monocular depth is relative, not metric. The synthesis is that class-agnostic segmentation supplies recall, geometry is the natural precision filter for it, and appearance OOD does not stand alone."

The evaluation uses the COCO validation set, designating seven classes as held-out and removing their leaves from the taxonomy: truck, bus (true super-category Vehicle) and horse, cow, sheep, elephant, bear (true super-category Living Being). The procedure yields n = 235 out-of-vocabulary ground-truth objects. Results: the flat/closed head emits a confident wrong specific label 100% of the time, with 37% of those in the wrong super-category (e.g., an animal named as a vehicle). The hierarchical layer emits zero confident wrong specific labels, recovers the correct super-category 26% of the time, flags 69% as honest UNKNOWN, and has 6% wrong super-branch errors, for a total of 94% safely handled. The authors are explicit: "This is the safety floor behaving as designed, a flagged unknown obstacle is preferable to a confident wrong guess, but it is not a claim of superior recognition accuracy. The layer's value is that it converts 100% confident-wrong into 0% confident-wrong at the price of frequent honest uncertainty. Corroboration on in-vocabulary objects shows 0% off-branch (categorical) errors with 24% calibrated abstention, versus a flat arg-max head's 53% off-branch errors on the same boxes."

The discussion emphasizes that hierarchical abstraction buys the elimination of confident categorical mistakes, and pays for it in specificity and abstention. The 69% abstention rate is called the principal limitation of the current system, attributed to the same scale/undersampling effect the feasibility study exposed. Future work targets reducing abstention via metric geometry from LiDAR or stereo, and motion cues, a dependable proposal front-end pairing class-agnostic recall with a geometric precision filter, and a hierarchical data engine that turns the layer's own open-world output into training data for a taxonomy-native detector.

Improvements for AI systems

Based on the paper, here are the specific improvements I can implement in an AI system, and what the improved system can do:

  1. Replace flat arg-max classification with mass-aggregated hierarchical abstraction
  • Instead of forcing a single leaf label, compute softmax over taxonomy leaves, aggregate probability mass upward to internal nodes, and descend only when a child concentrates ≥40% of local mass.

  • If no leaf exceeds an absolute similarity floor (0.20), or descent halts below the branch safety floor, emit UNKNOWN OBSTACLE.

  1. Add a class-agnostic proposal front-end (MobileSAM) unioned with closed-detector boxes
  • Run YOLO at low confidence with class-agnostic NMS, then union with MobileSAM masks so untrained objects still yield regions for classification.
  1. Integrate monocular depth as a geometric precision filter
  • For each proposal, compute foreground-vs-background separation and flat-vs-solid test from a depth map; suppress background regions and flag small/distant regions as not assessable rather than classifying them.
  1. Add a per-branch safety floor to prevent collapse into useless generic buckets
  • Define a floor per taxonomy branch (e.g., Vehicle, Living Being, Static Object); below the floor, report UNKNOWN OBSTACLE instead of a too-generic label.
  1. Use CLIP zero-shot scoring for every taxonomy node, not just leaves
  • Compute cosine similarity of region embedding to text prompts for all leaves; aggregate upward for internal nodes; this enables abstraction without retraining.
  1. Implement a leave-classes-out evaluation protocol for continuous monitoring
  • Hold out seven COCO classes (truck, bus, horse, cow, sheep, elephant, bear), prune their leaves, and stream validation images to measure safe-handling rate and abstention rate on out-of-vocabulary objects.

  • On out-of-vocabulary objects (e.g., horse-drawn carriage, road debris, livestock):

  • Never emits a confident wrong specific label (0% vs. flat head's 100%).

  • Avoids wrong-super-category errors (6% vs. flat head's 37%).

  • Safely handles 94% of such objects, either by correct super-category (26%) or honest UNKNOWN OBSTACLE (69%).

  • Provides a localized, actionable flag for planning rather than a silent misclassification.

  • On in-vocabulary objects:

  • Eliminates categorical (off-branch) errors (0% vs. flat head's 53%) while trading only 24% calibrated abstention, not a loss of correct answers.

  • On class-agnostic proposals:

  • Recovers regions the closed detector misses (148 additional regions per sample in the study).

  • Filters background/vegetation via the hierarchical UNKNOWN gate and depth-based flatness test, reducing false positives.

  • On ambiguous or undersampled crops:

  • Defaults to UNKNOWN OBSTACLE rather than forcing a guess, preserving safety.

  • For planning:

  • Provides a hierarchy of labels (e.g., some large living being ahead or unknown obstacle, localized here) that a planner can act on, versus a confident sedan attached to a standing horse.

  • For data collection:

  • Generates a hierarchically labeled corpus with a built-in review queue (the unknowns), which can be used to train a taxonomy-native detector, closing the loop from open-world handling to closed-set accuracy.

Cost to be explicit about: The system abstains on 69% of out-of-vocabulary objects, so it is safe but not specific. This is the intended safety trade-off, not a recognition-accuracy claim.

Abstract

A closed-set detector for autonomous driving must assign every object one of a fixed set of labels. On an object outside that set (a horse-drawn carriage, road debris, livestock on a rural road) it can only force a confident but wrong specific label or drop the object. Prior work in this series replaced the flat label set with a hierarchical taxonomy and a runtime abstraction rule, but evaluated it only on the boxes a closed detector already produces. This paper takes the layer open-world: we place taxonomic abstraction on top of class-agnostic region proposals so objects the closed detector never boxes can still be classified or flagged; we report a feasibility study of three open-world signals (class-agnostic segmentation, appearance-based out-of-distribution scoring, monocular depth) that shows why no single 2D cue suffices and how they compose; and we run the evaluation the earlier papers could not, a ground-truth leave-classes-out benchmark on real annotated objects. Holding out seven COCO classes and classifying their 235 ground-truth crops, a flat closed head emits a confident wrong specific label 100% of the time (37% of them in the wrong super-category, e.g. an animal named as a vehicle), whereas the hierarchical layer emits zero confident wrong specific labels and safely handles 94% of the objects (a correct super-category, or an explicit UNKNOWN OBSTACLE). We are explicit that this is a safety result, not a specificity one: the correct super-category is recovered only 26% of the time and the remaining 69% are conservatively flagged unknown. The contribution is an open-world perception layer that never makes a confident categorical mistake on an out-of-vocabulary object, together with an honest account of its cost.

Related papers