OpenBox: Annotate Any Bounding Boxes in 3D

arXiv:2512.01352 · cs.CV · Submitted 2025-12-01 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "OpenBox: Annotate Any Bounding Boxes in 3D".

Jane: OpenBox introduces a novel two-stage automatic annotation pipeline that leverages 2D vision foundation models to generate high-quality, open-vocabulary 3D bounding box annotations for vehicles, pedestrians, and cyclists without requiring self-training.

Tom: First, who's behind it and why it matters.

Title and authors: Tom: So, we're looking at the paper titled "OpenBox: Annotate Any Bounding Boxes in three dee," and the authors are Lee, Kim, Ryu, Musacchio, and Park from Seoul National University and POSTECH <ref:2512.01352#pg0,OpenBox: Annotate Any Bounding Boxes in 3D>. The title itself really tells you what it does: it lets you annotate any bounding boxes in three dee <ref:2512.01352#pg0,annotate any bounding boxes in 3D>.

Jane: That title suggests a very broad capability that's exactly what we need when we start thinking about open-vocabulary detection where we don't want to be stuck with just a few predefined object types.

Lu: I think the authors are smart because they are specifically targeting the limitations of existing methods, which is always a good sign for research; they aren't just tacking on a feature, they are fixing specific problems.

Meng: Fixing those problems usually means dealing with messy data or slow training cycles, so I’m curious how much effort goes into making sure this pipeline actually runs efficiently in real-world scenarios.

Lalam: It’s exciting to see research that focuses on the foundational structure of perception, moving beyond simple classification and into something that understands the physical presence of objects in space.

The paper's summary: Tom: The core idea of OpenBox is a two-stage pipeline that uses a vision foundation model to generate high-quality, open-vocabulary three dee bounding boxes for vehicles, pedestrians, and cyclists without needing self-training at all <ref:2512.01352#pg0>. It basically links cues from 2D images with the actual three dee point clouds through cross-modal instance alignment <ref:2512.01352#pg0>.

Jane: So it’s connecting what the vision model sees in a picture—like an instance ID or a class label—to where those objects are physically located in the LiDAR data, which is a really neat way to bridge that gap.

Lu: The summary highlights how they address the limitations of current approaches by associating instance-level cues with three dee point clouds and then generating adaptive bounding boxes based on the physical states of those instances <ref:2512.01352#pg0>. That adaptive part is what I find particularly compelling for handling different object behaviors.

Meng: So, if I understand correctly, it’s about taking a 2D semantic hint and using that to produce a physically informed three dee box, which should theoretically make the resulting annotations much more useful for downstream tasks <ref:2512.01352#pg0>.

Lalam: That focus on instance-level cues really helps us see objects as distinct entities rather than just points in a cloud, which is a fundamental shift in how we might process environmental data.

The paper's improvements: Tom: The paper points out several specific improvements they made, starting with cross-modal instance alignment where they map the 2D cues to the three dee point clouds <ref:2512.01352#pg0>. They also introduce context-aware refinement to clean up unprojected point clouds and an adaptive generation stage that chooses a fitting strategy based on the instance's physical state.

Jane: The context-aware refinement using majority voting within clustered regions sounds like a clever way they handled those tricky issues where LiDAR points might project onto background stuff like guardrails instead of the actual object.

Lu: And I’m particularly interested in how they categorize instances into rigid, dynamic rigid, and deformable types, because that decision directly influences the second stage of bounding box generation to use different fitting strategies for each.

Meng: So, they aren't just generating a single kind of box; they are tailoring the three dee output based on whether the object is static or if it’s moving or changing shape, which seems very practical for robotics <ref:2512.01352#pg0>.

Lalam: That ability to switch strategies based on physical state is powerful; it suggests that we can build systems that are inherently more robust because they aren't forced into a rigid geometric assumption for every single object they encounter.

Conclusion: Tom: To wrap up, OpenBox shows a two-stage automatic annotation pipeline that uses 2D vision foundation models to create open-vocabulary three dee bounding boxes by aligning instance cues with point clouds and generating adaptive boxes based on physical states <ref:2512.01352#pg0,a two-stage automatic annotation pipeline that>. It achieves high performance on datasets like the Waymo Open Dataset, even getting "seventy point four nine percent APthree dee for the vehicle class of the WOD at zero point five IoU."

Jane: So, to summarize, this work significantly reduces annotation costs and computational overhead by providing automated methods that respect an object's physical reality, which is a big step forward in making perception systems more scalable.

Lu: The implications here are huge because it opens up the door for developing general-purpose autonomous agents that can recognize things we haven't explicitly trained on before just from their descriptions.

Meng: From an engineering viewpoint, if this pipeline proves robust across different object types, it could significantly speed up deployment cycles for new perception features in real-world platforms.

Lalam: I think the ability to annotate arbitrary classes using semantic descriptions is really exciting because it means we can expand what our AI can perceive far beyond the initial training data constraints.

In-Jae Lee, Mungyeom Kim, Kwonyoung Ryu, Pierre Musacchio, Jaesik Park

Seoul National University · POSTECH

cs.CV

Submitted: 2025-12-01

Updated: 2026-10-02

Code: https://github.com/open-mmlab/mmdetection3d

Importance score: 89/100

The gist: OpenBox introduces a novel two-stage automatic annotation pipeline that leverages 2D vision foundation models to generate high-quality, open-vocabulary 3D bounding box annotations for vehicles,

Key concepts

Cross-modal Instance Alignment
This is the first stage where the system links information from a 2D image (like a vehicle's ID or class) to its corresponding points in a 3D point cloud. It establishes a mapping between these two different data types, ensuring that the correct 3D points are associated with the correct visual features, even when there are projection errors.
Context-aware Refinement
This step cleans up noisy unprojected LiDAR points by checking their relationship with surrounding objects. It uses majority voting within clusters and proximity ratios to decide which point clouds belong together, filtering out points that are incorrectly projected onto background elements like guardrails, thus improving the accuracy of the 3D structure.
Adaptive 3D Bounding Box Generation
The final stage classifies objects into rigid or deformable categories. Based on this classification, it generates tailored bounding boxes using class-specific size statistics. For example, static objects get one type of box, while dynamic ones get another, ensuring the resulting 3D box accurately reflects the object's physical movement and shape.
Physical State Categorization
Instances are sorted into three types: static rigid (like a parked car), dynamic rigid (like a moving car), and deformable (like a person or cloth). This categorization is crucial because the method uses different geometric refinement techniques for each type, such as surface-aware filtering for static objects versus closeness-to-edge fitting for deformable ones.

Terminology

Summary

OpenBox introduces a novel two-stage automatic annotation pipeline that leverages 2D vision foundation models to generate high-quality, open-vocabulary 3D bounding box annotations for vehicles, pedestrians, and cyclists without requiring self-training. This method addresses the limitations of existing approaches by associating instance-level cues from 2D images with corresponding 3D point clouds via cross-modal instance alignment and then generating adaptive bounding boxes based on the physical states of instances.

Cross-modal Instance Alignment

The first stage of OpenBox associates instance-level cues from 2D images processed by a vision foundation model with the corresponding 3D point clouds via cross-modal instance alignment. This process involves establishing a mapping between 2D image information, including instance IDs, class labels, and segmentation masks, and the corresponding 3D point cloud. The system obtains instance-level point clouds where each point contains 3D coordinate (x, y, z), semantic class, instance presence, and instance ID. To mitigate noise from imprecise mask boundaries due to calibration errors or projection issues caused by unprojecting masks directly into 3D space, the method employs an adaptive erosion proposed in [13], which erodes masks based on object size to eliminate boundary noise while preserving instance structure.

Context-aware Refinement

To enhance the quality of the unprojected instance-level point clouds, OpenBox applies a context-aware refinement step. This step addresses inaccurate unprojection where LiDAR points are projected onto background objects like guardrails or walls, which can lead to improperly scaled 3D bounding boxes. The refinement is performed by performing majority voting within clustered regions obtained from the ground-removed raw LiDAR point cloud using HDBSCAN [3]. Specifically, for each segment Rk, the method compares it with all instance-level point clouds Fi and computes bidirectional proximity-based inclusion ratios. A cluster Rk is retained as instance ID i if mutual overlap between the two clusters is sufficient, formulated by the condition: "Rk > α, f ∈ Fi dist(p, Fi) β."

Adaptive 3D Bounding Box Generation

The second stage categorizes instances by their physical states—static rigid, dynamic rigid, and deformable—and generates adaptive bounding boxes with class-specific size statistics. The process begins by densifying the refined instance-level LiDAR point clouds Fref by aggregating consecutive frames and using the PP score [43] to estimate the ephemerality of each point. Instances are then divided into three types: rigid and static F S ref, rigid and dynamic F D ref, and deformable F deform ref. The object type is determined using ChatGPT [27] based on the semantic class to distinguish between rigid and deformable objects.

Handling Static & Rigid Instances

For static instances, the method applies a surface-aware filtering method based on proximity voting over mesh vertices to suppress noise in the aggregated static point cloud Fref. A mesh surface S is reconstructed from the point cloud using the Signed Distance Function (SDF) [36]. For each vertex v ∈ S, points are retained only if foreground associations dominate, forming the refined surface Sref as defined by Eq. 4. The final bounding box is then refined via 3D-2D IoU alignment and visibility. This involves extracting the instance-level surface mesh Sins from S and comparing candidate boxes with projected boxes using instance ID, projecting them onto multiple views and time series images to select the box with the higher IoU.

Handling Dynamic & Deformable Instances

For dynamic rigid instances, OpenBox estimates orientation by aligning it with the direction of the object trajectory associated with 2D tracking IDs. The bounding box is then refined by computing dot products between outward surface normals and LiDAR ray directions at face centers. The box is extended only when the dot product between the ray and face normal is negative. For deformable instances, which exhibit articulated motion, geometry-based refinement is ineffective; instead, OpenBox generates bounding boxes from a single frame by tightly fitting the visible region using the closeness-to-edge algorithm [47], which provides robust representations without relying on rigid geometric assumptions.

Experimental Validation

Experiments are conducted on the Waymo Open Dataset (WOD), Lyft Level 5 Perception dataset, and nuScenes dataset. The method achieves high performance, for instance, achieving 70.49% AP3D for the vehicle class of the WOD [33] at 0.5 IoU. Qualitative results demonstrate that OpenBox produces high-quality and robust 3D annotations and enables the automatic annotation of open-vocabulary objects beyond predefined classes, such as strollers, fire hydrants, and dogs. The ablation study confirms that applying both point-level refinement modules yields the highest performance.

Limitations

The paper notes several limitations:

Improvements for AI systems

Based on the scientific paper OpenBox: Annotate Any Bounding Boxes in 3D, here are specific improvements that can be made to existing AI systems, and what those improved systems could achieve:


  1. Improve the efficiency and scalability of autonomous driving perception by enabling high-quality, open-vocabulary 3D object annotation without manual labor or iterative self-training.

  2. Enhance the accuracy of 3D object detection in complex urban environments by generating physically meaningful bounding boxes that account for an object's rigidity (static vs. dynamic) and motion state.

  3. Enable real-time, high-precision 3D localization and path planning systems by providing robust, instance-specific 3D bounding boxes that are refined through cross-modal alignment between 2D vision foundation models and LiDAR point clouds.

  4. Facilitate the development of general-purpose autonomous systems capable of detecting previously unseen or arbitrary object classes (open-vocabulary detection) based solely on semantic text descriptions, overcoming the limitations of fixed label spaces.

  5. Improve the robustness and accuracy of 3D perception systems in challenging conditions (e.g., occlusions, varying lighting) by incorporating context-aware refinement and surface-aware noise filtering on LiDAR point clouds before bounding box generation.

  6. Develop more reliable perception systems for dynamic objects (pedestrians, cyclists) by generating adaptive bounding boxes that account for their specific physical states (rigid vs. deformable) and motion characteristics, leading to less ghosting or distorted geometry in aggregated point clouds.

  7. Improve the generalization capability of 3D perception models by leveraging instance-level cues from powerful 2D vision foundation models (like Grounding DINO and SAM2) to guide the 3D annotation process, reducing reliance on purely geometric assumptions.

  8. Accelerate the iteration cycle for autonomous driving model training by providing high-quality, physical-state-aware annotations that require no self-training iterations for refinement, thereby significantly reducing computational overhead and time to deploy.

  9. Enable rapid prototyping of perception systems for novel scenarios (e.g., detecting specific street furniture like fire hydrants or strollers) by leveraging the open-vocabulary capabilities of OpenBox to annotate arbitrary classes directly from 2D prompts into 3D space.

Related papers