GEM-Occ: From Visual Geometry Evidence to Embodied Semantic Occupancy Memory
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.
Rosa: Today's paper: "GEM-Occ: From Visual Geometry Evidence to Embodied Semantic Occupancy Memory".
Dev: Semantic occupancy provides a structured spatial memory for embodied indoor agents by jointly representing occupied regions, observed free space, unknown areas, and object semantics.
Rosa: First, who's behind it and why it matters.
Paper summary: Rosa: So, we're diving into "GEM-Occ: From Visual Geometry Evidence to Embodied Semantic Occupancy Memory," which sounds like it’s tackling a big gap in indoor mapping. The thesis seems to be about creating a structured spatial memory that handles occupied regions, observed free space, unknown areas, and object semantics all at once for embodied agents <ref:2607.05543#pg0>. What claims are the authors making with this approach?
Dev: They're essentially aiming to solve the problem where existing indoor occupancy benchmarks often stick to single-view prediction or just room-level online perception, leaving long-horizon mapping across connected indoor spaces underexplored <ref:2607.05543#pg1>. The core claim of GEM-Occ is introducing HIOcc, a hierarchical benchmark that unifies ScanNet, ScanNet++, and Matterportthree dee under one sparse semantic occupancy format while keeping their original observation geometries intact <ref:2607.05543#pg0>. This setup allows for three different evaluation regimes: local prediction, room-level mapping, and building-level mapping across connected panoramic environments <ref:2607.05543#pg1>.
Taro: I'm interested in the structure of this memory because for autonomy researchers, the ability to handle context beyond immediate vicinity is crucial. If they're unifying these different data sources into a common format, it suggests a more robust way for an AI agent to build its long-term understanding of its surroundings <ref:2607.05543#pg1>. Does this structure actually support the kind of complex reasoning we need when things go wrong?
Rosa: Exactly, Taro; that hierarchy is key because it lets the system move from local geometry predictions up to a building-level graph <ref:2607.05543#pg1>. This isn't just a single map; it's a layered structure where you can query information at different scales of detail <ref:2607.05543#pg1>. It seems like they’re trying to build something that behaves more like persistent, meaningful spatial memory rather than just transient sensor readings.
Dev: From an engineering standpoint, the way they handle the evidence conversion is interesting; they propose GEM-Occ to convert transient visual evidence into persistent semantic Gaussian occupancy memory <ref:2607.05543#pg1>. They aren't sticking to pointmaps for map states; instead, they treat local geometry predictions as transient evidence and turn them into "semantic Gaussian occupancy evidence and free-space ray evidence" <ref:2607.05543#pg1>. That sounds like it could manage the flow of data better than a fixed pointmap structure would allow.
Taro: That idea of treating the local geometry predictions as transient evidence makes sense for handling dynamic environments where things move around quickly. But how does that transient evidence get stabilized into something persistent, especially when dealing with uncertainty in those predictions? I worry about the stability when the world misbehaves, like a door suddenly opening unexpectedly <ref:2607.05543#pg1>.
Paper summary: Rosa: The persistence comes from their causal updates; they use "visibility-and uncertainty-aware causal updates" to fuse this evidence <ref:2607.05543#pg1>. They organize the memory hierarchically into local caches, room-level submaps, and a building-level graph, which gives them multiple layers to check when things become unpredictable <ref:2607.05543#pg1>. This layered approach should help manage that uncertainty better than a single map representation would.
Dev: I'm looking at the math behind the fusion; they use a Mahalanobis-distance gate for matching incoming evidence primitives to existing memory primitives, followed by confidence-weighted fusion for updating the occupancy log-odds <ref:2607.05543#pg1>. This sounds like a sophisticated way to decide whether new information is relevant enough to update the map state or if it should be discarded due to low confidence. The latency implications here are something I need to watch closely, Rosa; how fast can this causal update actually run in real-time?
Taro: If the system can dynamically incorporate negative evidence from free-space ray evidence alongside positive occupied evidence, it should offer a better mechanism for understanding what *isn't* there yet <ref:2607.05543#pg1>. When an agent encounters an unknown area, having explicit information about the boundaries of that uncertainty helps in planning safer movements when things misbehave <ref:2607.05543#pg1>.
Rosa: That’s a big deal for deployment, Taro; if the system can reliably distinguish between truly unknown space and just areas we haven't seen yet, it offers a much more intelligent way to navigate and interact with an environment than just relying on simple occupancy predictions <ref:2607.05543#pg1>. It moves beyond just knowing where things are to understanding the spatial context of what is not there.
Dev: It’s important that this doesn't just work perfectly in a controlled simulation; we need to know how long this memory structure can maintain accuracy when faced with real-world sensor noise and drift <ref:2607.05543#pg1>. The authors mention they are testing local prediction, room-level online mapping, and building-level mapping, which suggests they're looking at the stability across different operational scales <ref:2607.05543#pg1>.
Taro: I’m hoping that when we look at the building-level mapping aspect of GEM-Occ, it shows how this memory persists over much longer periods than previous methods <ref:2607.05543#pg1>. That long-horizon capability is where true autonomy really gets tested when dealing with complex, interconnected spaces <ref:2607.05543#pg1>.
Rosa: Well, to wrap up this summary of "GEM-Occ: From Visual Geometry Evidence to Embodied Semantic Occupancy Memory," the paper proposes a unified framework for indoor spatial memory by integrating visual geometry evidence with semantic occupancy representation through HIOcc <ref:2607.05543#pg0>. The main contribution is building GEM-Occ, which uses Gaussian memory and causal updates to create persistent, hierarchical maps that handle local prediction, room-level online mapping, and building-level mapping across connected environments <ref:2607.05543#pg1>.
Paper summary: Dev: And the method relies on converting visual geometry predictions into semantic Gaussian primitives gti and using a causal fusion mechanism to update them based on observed evidence and ray evidence <ref:2607.05543#pg1>. The objective function involves several losses, including occupancy loss, semantic loss applied only to occupied voxels, geometry supervision for depth prediction, and a ray loss to manage free-space predictions <ref:2607.05543#pg1>.
Taro: The implications for autonomy are significant because this structure allows an agent to maintain a rich spatial understanding that connects local observations across entire buildings, which is vital when navigating complex, real-world indoor scenarios where the environment is constantly changing <ref:2607.05543#pg1>. It suggests a path toward more resilient long-term spatial reasoning for embodied AI systems.
Rosa: That's the big picture—moving from just seeing an immediate room to having a persistent, context-aware memory of the entire building structure <ref:2607.05543#pg1>. It moves the needle on how reliably robots can operate in unstructured indoor settings <ref:2607.05543#pg1>. We've heard about how this works conceptually, but we need to discuss what it actually means for deployment outside of a clean lab setting, and Rosa needs to ask that question next.
Dev: I agree; the transition from theoretical framework to reliable performance in a real-world loop rate context is where we have to focus our attention <ref:2607.05543#pg1>. The authors mention the evaluation regimes test room-level online mapping stability, which directly relates to how well this memory persists under sequential observations <ref:2607.05543#pg1>.
Taro: When we think about the real world, misbehavior is constant; an unexpected obstacle appearing or a sensor glitch causing a momentary false reading could derail a simple map system <ref:2607.05543#pg1>. I’m curious if this causal memory fusion mechanism can effectively recover from those sudden deviations without completely losing track of the larger context <ref:2607.05543#pg1>.
Rosa: That's the exact question I want to ask, Taro; does GEM-Occ have a mechanism that allows it to rapidly correct its understanding when the input evidence contradicts its existing memory state? It has these causal updates, so it should theoretically be able to incorporate that contradictory information in a weighted manner <ref:2607.05543#pg1>.
Dev: From my side, I need to know about the computational overhead of maintaining that hierarchical Gaussian memory structure; if the memory primitives get too complex or too numerous, we’re looking at unacceptable latency for real-time control loops <ref:2607.05543#pg1>. The authors don't give us a clear breakdown of the complexity involved in querying that building-level graph <ref:2607.05543#pg1>.
Paper summary: Taro: That complexity is a real concern for deploying this on resource-constrained hardware, Dev; if the system gets bogged down trying to reconcile all those layers constantly, it won't be useful when things get chaotic <ref:2607.05543#pg1>. The potential impact here could be creating agents that are much more robust against unexpected events because they have this layered understanding of space <ref:2607.05543#pg1>.
Rosa: So, we're looking at a system that aims to provide structured spatial memory through HIOcc and GEM-Occ, with the goal of enabling better long-horizon mapping across connected spaces <ref:2607.05543#pg1>. It’s an interesting step toward giving embodied agents a more coherent, persistent understanding of their indoor world <ref:2607.05543#pg1>. This paper certainly sets a good benchmark for what high-fidelity spatial memory looks like in this domain <ref:2607.05543#pg1>.
Dev: Indeed, the results show that GEM-Occ improves performance on local prediction, achieving sixty-one point three seven occupancy IoU and fifty-seven point seven six semantic mIoU on HIOcc when compared against GPOcc <ref:2607.05543#pg1>. Those comparative numbers suggest a tangible improvement in how the system handles local volumetric support <ref:2607.05543#pg1>.
Taro: Those quantitative improvements are encouraging, but I’m still focused on the practical implications for autonomy; if this framework can handle long-horizon mapping across connected buildings reliably, it opens up possibilities for autonomous agents to navigate complex environments without needing constant re-localization <ref:2607.05543#pg1>.
Rosa: That's exactly what I want to explore next; does the paper suggest any future work that focuses on extending this memory beyond indoor settings or testing its endurance in truly unstructured, unpredictable outdoor environments? The authors usually point toward where the next steps lie <ref:2607.05543#pg1>.
Dev: Looking at their stated limitations, they mention that existing benchmarks mainly focus on single-view prediction or room-level online perception, which is what GEM-Occ aims to address <ref:2607.05543#pg1>. However, they don't explicitly detail the long-term maintenance challenges for this persistent memory structure when faced with prolonged periods of sensor degradation or complete environmental changes <ref:2607.05543#pg1>.
Taro: So, it seems the current limitation lies in proving that this hierarchical Gaussian memory can maintain its integrity over very long operational horizons outside of controlled testing conditions <ref:2607.05543#pg1>. If they can solve that, it really validates the potential for more autonomous systems to operate reliably in complex indoor settings <ref:2607.05543#pg1>.
Rosa: That's a solid summary of where the work sits right now—a powerful framework for unified indoor mapping, but with clear avenues remaining to prove its robustness in long-term, unconstrained operational scenarios <ref:2607.05543#pg1>. It definitely gives us a lot to chew on as field roboticists <ref:2607.05543#pg1>.
Conclusion: Rosa: So, we've seen how GEM-Occ builds this structured spatial memory using Gaussian primitives and causal updates to handle everything from local geometry to building maps across connected spaces.
Dev: Yeah, that framework sounds incredibly dense, but the core idea is taking raw visual data and turning it into a persistent map structure that respects both what we see and what we infer about free space.
Taro: I’m really thinking about how this memory handles uncertainty when the agent encounters something unexpected or when sensor readings are noisy in a real-world setting.
Rosa: Exactly, Taro; my main question is whether this level of structural understanding translates well to the messiness of an actual indoor environment over a long operational period.
Dev: From my angle, I’m worried about the computational load; maintaining that hierarchical Gaussian memory structure must be taxing on loop rates and we need to know how it handles potential failure modes in real-time.
Taro: That’s a fair point, Dev; if the system gets bogged down trying to reconcile all those layers constantly, it won't be useful when things get chaotic or unpredictable.
Rosa: So, looking at the title 'GEM-Occ: From Visual Geometry Evidence to Embodied Semantic Occupancy Memory', what does that really mean in plain terms for us?
Dev: It means they’re connecting the visual data—the geometry—directly to a meaningful semantic map, which is a big step past just tracking where objects are located.
Taro: And it moves beyond simple occupancy maps by incorporating both observed free space and unknown areas into one coherent structure.
Rosa: That sounds like it gives an agent a much richer context than before, allowing it to reason about its surroundings on multiple spatial levels simultaneously.
Dev: I see the implication for deployment is that if we can nail the loop rate, this could enable agents to navigate complex indoor spaces with a level of long-term awareness we haven't seen yet.
Taro: The real impact could be in autonomous systems that need to operate reliably in environments where things aren't perfectly structured or predictable.
Rosa: So, we’re talking about a system that moves beyond just mapping rooms to understanding the spatial relationships across entire buildings with persistent memory.
Dev: That persistence is key; if it can maintain that accuracy over extended periods, it opens up possibilities for agents to function in more complex, long-term tasks.
Taro: I’m still keen on how robust this memory proves itself when faced with prolonged periods of sensor degradation or major environmental changes outside of the controlled lab setting.
Rosa: That’s exactly where we need to keep our focus moving forward; proving that endurance is what separates a theoretical success from a practical tool for field robotics.
Hu Zhu, Bohan Li, Xianda Guo, Hongsi Liu, Baorui Peng, Mingqi Yuan
The Hong Kong Polytechnic University · Eastern Institute of Technology Department of Computer Science and Engineering, Shanghai Jiao Tong University
cs.RO, cs.CV
Submitted: 2026-07-06
Updated: 2026-10-05
Comments: Project page: https://zhuhu00.top/GEM-Occ/
Code: https://github.com/OpenDriveLab/OpenScene
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 88/100
The gist: Semantic occupancy provides a structured spatial memory for embodied indoor agents by jointly representing occupied regions, observed free space, unknown areas, and object semantics.
Key concepts
- Semantic Gaussian Evidence
- This is the output from an encoder that predicts local geometry, image features, semantic labels, and confidence for each observation. It is modeled as a Gaussian primitive that captures the spatial distribution of observed data, where the covariance shape follows the viewing ray to accurately represent uncertainty.
- Free-Space Ray Evidence
- This evidence accumulates information about areas not yet hit by any surface during an observation. It helps distinguish between occupied and unknown regions by tracking rays that pass through empty space, which is crucial for accurate mapping in complex environments.
- Hierarchical Gaussian Memory
- The persistent map is organized into a hierarchy: local caches, room-level submaps, and a building-level graph. This structure allows the system to store and query spatial information at different scales, enabling flexible mapping across various levels of detail.
- Causal Memory Fusion
- This mechanism updates the persistent memory when new evidence arrives. It uses a spatial search and a distance gate to match new evidence with existing memories, then fuses them using confidence-weighted formulas to refine both the occupancy state and semantic labels.
Terminology
Summary
Semantic occupancy provides a structured spatial memory for embodied indoor agents by jointly representing occupied regions, observed free space, unknown areas, and object semantics.
How it works
-
HIOcc (Hierarchical Indoor Occupancy benchmark) is introduced to unify ScanNet, ScanNet++, and Matterport3D under a common sparse semantic occupancy format while preserving their native observation geometries (perspective RGB-D frames and pano-centric observation groups). This benchmark supports three complementary evaluation regimes: local semantic occupancy prediction, room-level online occupancy mapping, and building-level mapping across connected panoramic environments.
-
GEM-Occ (Gaussian Evidence Memory framework) is proposed to convert transient visual evidence into persistent semantic Gaussian occupancy memory. Instead of using pointmaps as persistent map states, GEM-Occ treats local visual geometry predictions as transient evidence and converts them into
semantic Gaussian occupancy evidence and free-space ray evidence.
-
This fusion is achieved through
visibility-and uncertainty-aware causal updates.
The memory is organized hierarchically intolocal caches, room-level submaps, and a building-level graph,
allowing it to be queried at any time throughGaussian-to-occupancy splatting.
Key Components of GEM-Occ
(The paper enumerates the following components and mechanisms in detail)
-
Semantic Gaussian Evidence: For each observation, an encoder predicts local geometry, image features, semantic logits, and confidence as (Dt, Ft, St, Qt). This is instantiated as a semantic Gaussian primitive gti = (µti, Σti, αti, pti, ηti), where the covariance follows the viewing ray: Σta1 = Rat1idiag(σ2⊥1, σ2⊥1, σ2∥1, i)(Rat1)⊤.
-
Free-Space Ray Evidence: Samples before the first surface hit are accumulated as free-space ray evidence, while regions behind the first hit remain unknown unless observed later.
-
Hierarchical Gaussian Memory: The persistent map is represented by Mt = (G occt, Rfreet), where G occt is semantic Gaussian occupancy memory and Rfreet is a sparse free-space ray cache. Each memory primitive Gtⱼ stores geometry, occupancy log-odds (ltⱼ), accumulated weight (w tⱼ), and observation count (n tⱼ).
-
Causal Memory Fusion: Incoming evidence primitives are matched to existing memory primitives using a local spatial search followed by a Mahalanobis-distance gate. If matched, the primitive is updated via confidence-weighted fusion: µtⱼ = (λwt−1j µt−1j + ηtiµti w¯tⱼ) / w¯tⱼ, and the occupancy log-odds are updated by jointly incorporating positive occupied evidence and negative free-space evidence: ltⱼ = lt−1j + ηti∆locc − βt−1j∆lfree.
Training Objective and Memory Maintenance
(The training objective is defined as follows)
-
The objective function is formulated as L = λoccLocc + λsemLsem + λgeoLgeo + λrayLray, where:
-
Occupancy loss (Locc) supervises occupied and free states under the valid observation mask.
-
Semantic loss (Lsem) is applied only to occupied voxels.
-
Geometry supervision (Lgeo) supervises the predicted depth or point geometry when available.
-
Ray loss (Lray) penalizes occupied predictions in observed free space and improves the separation between free and unknown regions.
Evaluation Regimes
(GEM-Occ is evaluated across three settings)
-
Local prediction: Predict semantic occupancy from one posed RGB observation using the local cache.
-
Room-level online mapping: Use a causal posed-observation sequence to evaluate the map after online fusion, testing stability over sequential observations.
-
Building-level mapping: Use Matterport3D splits where each sample is a panorama with rectified sub-cameras to test long-horizon mapping across connected panoramic environments.
Results Summary
(Experiments show that GEM-Occ improves performance over baselines)
-
Local prediction: GEM-Occ achieves the best overall performance on HIOcc, reaching 61.37 occupancy IoU and 57.76 semantic mIoU, outperforming GPOcc by providing stronger local volumetric support than directly accumulating pointmap-derived or frame-level Gaussian predictions.
-
Room-level mapping: GEM-Occ improves IoU from 52.94 to 56.79 and mIoU from 44.
Improvements for AI systems
Here are the specific improvements that can be made to AI systems by implementing the proposed HIOcc and GEM-Occ framework, and what those improved systems will be able to do:
-
The ability of embodied agents (robots) to perform robust, long-horizon navigation and exploration in complex, connected indoor environments.
-
The capability for real-time
online
semantic mapping during exploration, maintaining a stable map representation across sequential observations without catastrophic forgetting or drift. -
The capacity for accurate reasoning about free space versus unknown space during movement (e.g., distinguishing between an empty hallway and a region behind an occluding object).
-
The ability to generate high-fidelity, scalable 3D semantic occupancy maps that can be queried at any arbitrary location (local, room-level, or building-level) for downstream tasks.
-
The enhancement of semantic understanding by explicitly encoding
semantic uncertainty
andvisibility history
into the persistent map state, leading to more reliable semantic classifications.
This is achieved through the following technical improvements:
-
By using the hierarchical benchmark (HIOcc), AI systems will be trained on data that covers a full range of indoor scales—from single-frame local prediction to building-level panoramic mapping—ensuring generalizability across different sensing modalities (RGB-D, perspective, and panoramic).
-
By implementing the Gaussian Evidence Memory framework (GEM-Occ), AI systems transition from treating occupancy as a static voxel grid or per-frame prediction to maintaining a dynamic, persistent map state composed of semantic Gaussians and free-space ray caches. This allows for causal updates based on observation sequences, enabling continuous refinement of the map over time.
-
By fusing local visual geometry predictions (depth/pointmaps) with explicit free-space ray evidence through a visibility-aware causal update rule, AI systems will accurately model the distinction between observed empty space and unknown regions, which is critical for safe navigation and planning.
-
By organizing the memory hierarchically (local caches, room submaps, building graphs) and using Gaussian-to-occupancy splatting for querying, AI systems gain the ability to perform efficient
semantic queries
to determine occupancy at any point in space with a quantifiable probability of being occupied or free. -
The system will be able to leverage semantic distributions (from predicted logits) and confidence metrics to produce not just an occupancy map, but also a probabilistic semantic distribution over the scene, allowing for more nuanced downstream tasks like semantic-goal navigation or object search based on learned uncertainty.
Abstract
Embodied agents exploring indoor environments require reliable semantic occupancy memory that persists across observations and revisits. Building such memory is challenging because each observation provides incomplete and uncertain geometric and semantic evidence. We introduce GEM-Occ, a Gaussian Evidence Memory framework that consolidates evidence accumulated over time into persistent semantic occupancy memory. Local predictions are converted into occupied semantic Gaussians and free-space ray evidence. Confidence- and visibility-aware causal updates integrate supporting observations, suppress occupancy contradicted by observed free space, and preserve previously observed structures through occlusion. A hierarchical memory organization supports continued mapping and efficient queries across connected indoor spaces. To evaluate this capability, we introduce HIOcc, a unified benchmark for embodied semantic occupancy memory. HIOcc establishes a shared semantic label space and evaluation framework spanning local prediction, room-level online mapping, and building-level mapping, while accommodating perspective and panoramic observations. Experiments on HIOcc demonstrate that GEM-Occ outperforms existing methods, enabling accurate semantic occupancy prediction and consistent online mapping across spatial scales with efficient memory usage and fast occupancy queries.
Sources
- Scene as Occupancy
- Bridging Stereo Geometry and BEV Representation with Reliable Mutual Interaction for Semantic Scene Completion
- RoboOcc: Enhancing the Geometric and Semantic Scene Understanding for Robots
- SGR-OCC: Evolving Monocular Priors for Embodied 3D Occupancy Prediction via Soft-Gating Lifting and Semantic-Adaptive Geometric Refinement
- Humanoid Occupancy: Enabling A Generalized Multimodal Occupancy Perception System on Humanoid Robots
- Bridging 3D Gaussians and Semantic Occupancy for Comprehensive Open-Vocabulary Scene Understanding from Unposed Images
- From Articulated Kinematics to Routed Visual Control for Action-Conditioned Surgical Video Generation
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving