Seeing Fast and Slow: Bimodal 3D Scene Graphs for Open-set Tasks

arXiv:2605.31067 · cs.RO · Submitted 2026-05-29 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: Today's paper: "Seeing Fast and Slow".

Dev: Open-set task execution can significantly benefit from seamlessly switching between coarse and fine scene representations depending on the context and the evolving information as the robot explores the environment.

Rosa: First, who's behind it and why it matters.

Paper summary: Rosa: So, wrapping up the discussion on "Seeing Fast and Slow: Bimodal three dee Scene Graphs for Open-set Tasks," it really boils down to how this bimodal approach allows a robot to operate efficiently by using a fast, closed-set representation initially and only engaging the slower, open-set generation when task relevance demands finer detail <ref:2605.31067#pg0,Seeing Fast and Slow: Bimodal 3D Scene Graphs for Open-set Tasks>.

Dev: The authors' title itself captures the essence of their contribution perfectly because they are tackling the need for both speed and accuracy simultaneously in open-set scenarios.

Taro: I think this framework has implications for deploying autonomous systems in environments that we can't fully model beforehand, making it much more practical for real-world exploration tasks.

Rosa: Exactly, and the fact that they demonstrate significant speed improvements over existing methods is what makes this work relevant for integrating scene graph generation directly into the task execution loop.

Dev: That speed difference is critical because it allows us to move closer to true real-time performance where decision-making and perception happen concurrently without major delays.

Taro: Ultimately, this suggests that future autonomy will rely less on building massive, static models of the world and more on intelligent systems that can intelligently adapt their level of scene detail on the fly.

Rosa: That's a big picture idea; we're moving toward systems that are inherently adaptive rather than just perfectly pre-programmed for specific scenarios.

Dev: It feels like a step toward making robots truly versatile explorers, capable of handling unexpected situations without getting bogged down in computational overhead.

Conclusion: Rosa: So, to wrap up this discussion on "Seeing Fast and Slow: Bimodal three dee Scene Graphs for Open-set Tasks," we've seen how this method lets robots switch between a quick overview and a detailed look at objects when they encounter something new.

Dev: I agree, the core idea is using two different speeds to handle the complexity of an open-set environment without getting completely bogged down by processing power.

Taro: I think what's really interesting is how this framework handles situations where the robot sees something it hasn't encountered before and needs to figure out what it actually is.

Rosa: Exactly, and looking at the authors, they seem to have put together a really neat system that tackles this dual need for speed and accuracy in a practical way.

Dev: From an engineering standpoint, the fact that they manage this switching mechanism efficiently is key; I'm curious about how stable that transition is under heavy load.

Taro: It's not just about the technical mechanism itself, though; it opens up possibilities for autonomy in much more unpredictable real-world settings where we can’t pre-model everything.

Rosa: That’s a big picture point, Taro; it suggests systems can be much more adaptable to novel situations than they are right now.

Dev: And I'm still focused on the practical reality of deployment; how long can this system reliably operate in a dynamic environment before those transition modes start introducing unacceptable latency?

Taro: That's where the limitations become apparent, though, because even with this bimodal approach, there are still moments where the system might misinterpret an object's function entirely if the coarse model is too misleading.

Rosa: That's a fair caveat; we can’t just assume perfect performance across all possible novel objects.

Dev: So, while it shows great potential for speed and flexibility, we need to see robust testing that proves this works reliably outside of a controlled lab setting over extended periods.

Singapore University of Technology and Design (SUTD)

cs.RO

Submitted: 2026-05-29

Updated: 2026-10-05

Comments: Submission has been approved by the funding agency

Code: https://github.com/luca-medeiros/lang-segment-anything

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 77/100

The gist: Open-set task execution can significantly benefit from seamlessly switching between coarse and fine scene representations depending on the context and the evolving information as the robot explores

Key concepts

Fast Mode (System-1)
This default mode quickly generates a coarse 3D scene graph using a closed vocabulary. It employs a novel prismatic representation of 3D objects to abstract point clouds into simpler shapes, drastically reducing computation and energy costs while providing enough information for general object reasoning.
Slow Mode (System-2)
This mode is activated intermittently to generate open-vocabulary scene graphs for task-relevant objects. It uses visual reasoning, specifically a VLM, to perform open set segmentation on objects detected by the planner, allowing the system to identify and integrate novel items into the scene graph.
Prismatic 3D Object Representation
This representation converts object masks into polygonal prisms. By projecting points onto the mask contour and constructing a prism shape, it creates a simplified, cuboid-like abstraction of 3D objects. This is computationally efficient and suitable for coarse scene graph construction.
Vicinity Based Object Relation
This technique establishes hierarchical relationships between 3D objects by first creating visibility matrices and pruning edges based on distance thresholds. It then orders nodes by an 'absolute status' to partition the scene graph into blocks, leading to a weighted quotient visibility graph.

Terminology

Summary

Open-set task execution can significantly benefit from seamlessly switching between coarse and fine scene representations depending on the context and the evolving information as the robot explores the environment. The gist: BiMoSG proposes a bimodal 3D scene graph generation approach that employs a fast mode for efficient coarse representation and switches to a slow mode for generating finer open-vocabulary 3D scene graphs of task-relevant objects, demonstrating significant speed improvements over state-of-the-art methods.

How it works

BiMoSG utilizes two distinct modes for 3D scene graph generation: a default fast or System-1-like mode and an intermittent slow or System-2-like mode. The fast mode is employed by default to efficiently generate a coarse 3D scene graph. This involves using a closed vocabulary setup to produce a coarse approximation of the environment objects, which is sufficient for reasoning about the general functionality of objects and spaces. To achieve this, the approach proposes a novel prismatic representation of 3D objects which provides a coarse representation, constructing an abstraction of 3D point clouds into prism representations to greatly reduce computation cost. Furthermore, it introduces a novel vicinity based approach to establish the hierarchical relationship between 3D objects in the scene graph that is built upon our prismatic representation. Crucially, the default fast mode does not require the use of large visual foundation models to construct scene graphs, drastically reducing the computation and energy cost of generating the scene graph.

How it works (Cont.)

The system can then intermittently switch to a second slow or System-2-like mode. This slow mode is used for generating an open-vocabulary scene graph, which can then be used to locate and retrieve objects with visual reasoning. The process involves:

  1. Using the closed segmentation using Mask2Former by default to generate a coarse scene graph.

  2. If the LLM Planners detects objects that matches one of the triggered object labels, it can trigger the ”slow” open set object identification module using VLM to perform open set segmentation.

  3. The identified open-set objects are then appended to the existing scene graph, combining closed-set and open-set segmentation models.

How it works (Cont.)

The methodology involves several key components:

A) Closed Set Segmentation: This uses Mask2Former [2], which unifies panoptic, instance, and semantic segmentation. It is specifically employed to provide both instance-level and semantic-level information.

B) Prismatic 3D Object Representation: This representation is constructed from object masks by taking the contour of the mask, projecting dense points on the contour line into world frame to obtain a 3D world-frame contour ring C, denoising it, and then constructing a polygonal prism O. This representation is chosen because modern furniture is designed to be packable inside cuboid or prismatic shaped containers, making it suitable as a coarse representation.

C) Vicinity Based Object Relation: This approach computes vicinity based relations between 3D objects by first constructing an n × n symmetric visibility matrix VM. It then iteratively prunes edges on the graph if the visible and effective distance dv between two objects is less than a threshold, defined as dv(i, j) = min(d(r i h + r j h), d(r i v + r j v)). It then orders nodes based on an absolute status SAi to rank their importance over the entire graph, leading to a hierarchical partitioning of the scene graph SG into Vicinity Vk blocks. Finally, a weighted quotient visibility graph (GV) is constructed where each vicinity Vk forms a supernode.

How it works (Cont.)

D) Planner: The LLM Planner acts as the decision-maker in a three-phase loop:

Phase 1 – Task Decomposition: The LLM analyzes the task and generates a set of trigger labels and corresponding object IDs from closed-set segmentation classes. These triggers identify objects likely to be relevant, including functionally related items not explicitly mentioned in the instruction.

Phase 2 – Scene Monitoring for Open-Set Activation: The planner monitors the vicinity graphs for objects or vicinities that might be of interest. When a match is detected, it invokes the InspectAt action on the corresponding object ID and prompts the VLM to refine the object’s class label.

Phase 3 – Object Relevance Confirmation: Once the updated scene graph is received from the VLM, the planner uses an LLM to confirm whether the inspected object is truly relevant to the task, ensuring alignment with high-level goals.

How it works (Cont.)

E) Bimodal Scene Graph: The system employs two ways of doing open set object detection: a hybrid mode and a switching mode.

Improvements for AI systems

Here are specific improvements that can be made to existing AI systems, derived from the BiMoSG framework:


The core innovation lies in developing a computationally efficient, context-aware perception pipeline for open-set tasks by leveraging a bimodal scene graph approach. The resulting system is specifically optimized for real-time deployment on resource-constrained edge devices (like mobile robots).

Here are the specific improvements and capabilities:

  1. Transition to Efficient, Hierarchical Scene Graph Generation (The Fast Mode)

Improve existing systems that rely solely on large Visual Foundation Models (LVLMs) for 3D scene graph construction by implementing a dual-mode generation strategy.

  • The system should default to a Fast Mode using a closed vocabulary setup and novel prismatic volumetric representations. This mode generates coarse, computationally cheap approximations of the environment objects and their spatial relationships, allowing for rapid initial reasoning about general object function (System 1 thinking).

  • The system can then dynamically switch to a Slow Mode only when encountering regions requiring high granularity, leveraging Language Models (LLMs) and Vision-Language Models (VLMs) to generate fine-grained open-vocabulary scene graphs for task execution refinement.

  1. Enhanced Real-Time Object Perception and Tracking

Improve real-time object mapping and tracking capabilities in dynamic environments.

  • Implement a novel Vicinity Based Approach that constructs hierarchical spatial relationships (Vicinities, Zones) between objects based on their relative size (Girth) and proximity, rather than just raw geometric distance. This allows the robot to prioritize relevant objects for interaction immediately.

  • Utilize volumetric object representations derived from mask contours instead of dense point clouds for initial scene graph generation. This offers computational advantages in projection, denoising, and merging/deduplication (achieving near O(n+m) complexity compared to O(nm) for point clouds), significantly speeding up the mapping process on mobile platforms.

  1. Context-Aware Task Planning and Open-Set Reasoning

Improve the decision-making layer of AI agents performing open-set tasks.

  • Integrate an LLM Planner that explicitly manages the mode switching logic (Closed Set vs. Open Set). The system should use closed-set detection for initial object grounding (Phase 1) and only invoke computationally expensive VLM refinement/open-set segmentation (Phase 2) when a potential candidate is identified, drastically reducing energy cost and latency.

  • In the Switching Mode, the system should utilize multimodal reasoning (textual vicinity graphs from closed sets combined with VLM descriptive captions for open sets) to ensure that object retrieval decisions are semantically grounded in the task context, minimizing false positives.

  1. Robust Semantic Grounding and Verification

Improve the accuracy and reliability of object identification in novel environments.

  • Implement a rigorous LLM-based confirmation step (Phase 3) where the agent queries an LLM with both the task description and detected object to verify relevance, preventing the system from acting on visually similar but contextually irrelevant objects.

  • Use CLIP embeddings for closed-set object representations to facilitate efficient volume merging of similar classes, improving deduplication in sparse input scenarios (e.g., fast movement).

The improved AI system can perform the following specific actions:

  1. It can execute open-set tasks (like find a suspicious object) in real-time on mobile robots without needing extensive prior mapping data.

  2. It can prioritize exploration by efficiently generating a coarse map of known objects first, significantly reducing initial processing time (up to 3x faster than state-of-the-art baselines).

  3. It can dynamically switch its perception granularity: using fast, low-cost closed-set detection for general navigation and switching to slow, high-fidelity open-set VLM reasoning only when a task demands precise identification of novel objects.

  4. It can perform complex reasoning tasks (e.g., where should I put this item?) by integrating spatial proximity analysis (Vicinity Graphs) with semantic understanding derived from both closed and open vocabularies, leading to more flexible and contextually relevant object interactions than purely retrieval-based systems.

Abstract

Open-set task execution can significantly benefit from seamlessly switching between coarse and fine scene representations depending on the context and the evolving information as the robot explores the environment. For example, it is often sufficient to start with a coarse scene representation initially and only employ a finer, more granular scene representation when the robot encounters regions which are likely to contain the task relevant objects. Hence, in this work, we propose BiMoSG, a bimodal 3D scene graph generation approach for open-set tasks. BiMoSG employs a "fast" mode by default to efficiently generate a coarse 3D scene graph and can switch to a "slow" mode for generating a finer open vocabulary 3D scene graph of task relevant objects. We demonstrate that our proposed 3D scene graph generation approach is significantly faster than the open-source state-of-the-art approaches. This allows us to integrate the scene graph generation process with task execution for real-time deployment.

Sources

Related papers