M squared-Occ: Resilient 3D Semantic Occupancy Prediction for Autonomous Driving with Incomplete Camera Inputs

arXiv:2603.09737 · cs.CV, cs.RO, eess.IV · Submitted 2026-03-10 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "M squared-Occ: Resilient 3D Semantic Occupancy Prediction for Autonomous Driving with Incomplete Camera Inputs".

Jane: Semantic occupancy prediction enables dense 3D geometric and semantic understanding for autonomous driving, but existing camera-based approaches implicitly assume complete surround-view observations,

Tom: First, who's behind it and why it matters.

Title and authors: Tom: Let's talk about the title and who came up with this. The paper is "M squared-Occ: Resilient three dee Semantic Occupancy Prediction for Autonomous Driving with Incomplete Camera Inputs," and the authors are Kaixin Lin, Kunyu Peng, Di Wen, Yufan Chen, Ruiping Liu, and Kailun Yang <ref:2603.09737#pg0,Occ: Resilient 3D Semantic Occupancy Prediction for Autonomous Driving with Incomplete Camera>.

Jane: It’s interesting to see the focus on resilience right in the title; they aren't just trying to make a model that works perfectly under ideal conditions. The paper is specifically designed for those tricky real-world situations where the sensor data isn't complete.

Lu: The team has put together a framework that builds on existing ideas, drawing inspiration from how humans infer unseen areas from context and memory, which is a really interesting angle for an AI system to consider.

Meng: I’ve seen some work on related topics like M-BEV or UniBEV, and this paper seems to be pushing the boundaries by focusing specifically on dense three dee semantic occupancy prediction under incomplete surround-view cameras <ref:2603.09737#pg0,semantic occupancy prediction under incomplete>. That sounds like a lot of data to handle efficiently.

Lalam: It’s exciting because it means we can build AI that is less fragile when things go wrong in the field, which is exactly what autonomous systems need to be trustworthy for everyone using them on the road.

The paper's summary: Tom: So, let's get into what they actually did. The M squared-Occ framework introduces a method that handles missing views by focusing on two main parts: the Multi-view Masked Reconstruction module for geometry and a Feature Memory Module for semantics.

Jane: That’s right, Tom. To put it simply, the MMR module uses the overlap between neighboring cameras to recover what's missing in the feature space directly, without having to generate raw pixels like some other methods might do. It’s essentially a "soft repair" mechanism operating within the features themselves.

Lu: And on top of that geometric recovery, they use a separate Feature Memory Module that stores class-level semantic prototypes. This module then goes back and refines those ambiguous three dee voxel features by adding corrections based on these learned global semantic priors <ref:2603.09737#pg0>.

Meng: So, one part handles the physical structure reconstruction through neighboring views, and the other part uses stored knowledge to stabilize what those reconstructed voxels actually represent semantically. That separation of concerns sounds very intentional for stability.

Lalam: That’s a really smart way to approach it; you get structural recovery from local context and semantic grounding from global knowledge, which should make the resulting three dee map much more reliable even when parts are missing <ref:2603.09737#pg0>.

The paper's improvements: Tom: The paper details some specific technical improvements they made to their existing occupancy prediction models. They propose the Multi-view Masked Reconstruction module, which uses a transformer-based decoder and a learnable mask token to synthesize a structural prior when an input view is masked.

Jane: And they quantify this using an "MMR Loss Function" calculated only on the set of masked view indices, which makes sure that the features they reconstruct are actually aligned with what the ground truth should look like in terms of semantics.

Lu: The FMM part is also quite detailed; they use two strategies for storing semantic prototypes: a single global centroid updated by momentum averaging, or a Multi-Proto strategy that learns multiple sub-prototypes per class and retrieves them based on cosine similarity.

Meng: I noticed the ablation study shows that MMR alone improves the IoU to twenty-eight point one nine percent, which confirms its role in restoring large-scale structures, but then they find that the FMM adds another layer of refinement to that result <ref:2603.09737#pg2>.

Lalam: It’s interesting how they found that under missing-view conditions, using a single global centroid seems more robust than having multiple subcentroids when the observations are really corrupted; it suggests simplicity can be powerful in certain failure modes.

Conclusion: Tom: So, to wrap things up, we've seen how M squared-Occ combines feature-level reconstruction from overlapping views with semantic regularization from a learnable memory bank to tackle incomplete camera inputs. The results show it significantly improves performance when views are missing compared to the baseline models they tested.

Jane: It really highlights that for autonomous systems, robustness against sensor failure isn't just about making one part work better; it’s about creating a system that can infer structure and maintain semantic meaning even when parts of the input data are gone.

Lu: The implication here is that we are moving toward AI systems that can exhibit a kind of contextual reasoning, where they don't just rely on what's immediately visible but use learned knowledge to fill in the gaps intelligently.

Meng: From a practical standpoint, this means we could deploy these systems in environments with less consistent sensor data, which is a huge step toward making autonomous driving truly reliable outside of perfect conditions.

Lalam: This work on M squared-Occ shows us that AI can be incredibly versatile when dealing with incomplete information; it’s about building systems that can hallucinate the necessary context to maintain safety and coherence across all scenarios.

Kaixin Lin, Kunyu Peng, Di Wen, Yufan Chen, Ruiping Liu, Kailun Yang

Hunan University

cs.CV, cs.RO, eess.IV

Submitted: 2026-03-10

Updated: 2026-10-05

Code: https://github.com/qixi7up/M2-Occ

Importance score: 89/100

The gist: Semantic occupancy prediction enables dense 3D geometric and semantic understanding for autonomous driving, but existing camera-based approaches implicitly assume complete surround-view observations,

Key concepts

Multi-view Masked Reconstruction (MMR)
This module recovers missing geometric data by treating the camera layout as a graph. When a view is missing, it stitches together features from adjacent cameras using overlapping regions and a learnable token. A transformer decoder then refines these stitched features to reconstruct the original structure, ensuring geometric continuity across occluded areas.
Feature Memory Module (FMM)
This component stores semantic knowledge in a memory bank of class-level prototypes. It uses either a single global centroid or multiple sub-prototypes per class. These prototypes are retrieved based on similarity to the current input features, and their weighted information is added back to the 3D volume as a residual correction to resolve semantic ambiguities.
MMR Loss Function (LMMR)
This loss function is specifically designed for training the geometric reconstruction part of M2-Occ. It only calculates errors on the features corresponding to the missing views. This forces the model to focus its learning on accurately reconstructing the structure where data is absent, ensuring that recovered geometry remains semantically relevant.

Terminology

Summary

Semantic occupancy prediction enables dense 3D geometric and semantic understanding for autonomous driving, but existing camera-based approaches implicitly assume complete surround-view observations, an assumption that rarely holds in real-world deployment due to occlusion, hardware malfunction, or communication failures.

The gist

M2-Occ introduces a framework designed to preserve geometric structure and semantic coherence when views are missing by leveraging spatial overlap among neighboring cameras for feature recovery and using a learnable memory bank to refine ambiguous voxel features with global priors.

How it works

The proposed M2-Occ framework is built on two complementary pillars: the Multi-view Masked Reconstruction (MMR) module for geometric recovery and the Feature Memory Module (FMM) for semantic regularization.

  1. The overall architecture follows a standard 2D-to-3D paradigm: input images are processed by a shared image backbone with a Feature Pyramid Network (FPN) to extract multi-scale 2D features, which are then lifted into a unified 3D volume space, and finally processed by a 3D occupancy head to predict the voxel-wise semantic labels.

  2. The Multi-view Masked Reconstruction (MMR) module addresses missing views by modeling the physical layout as a cyclic graph to exploit spatial redundancy. When a view is masked, it performs feature cropping and splicing operation on the overlapping boundary regions of neighboring cameras, concatenating them with a learnable mask token to synthesize a structural prior. This synthesized feature is then fed into a lightweight transformer decoder that refines the features to approximate the original unmasked representation using the formula:

ˆfi = D(fref + ppos).

  1. The MMR module utilizes an MMR Loss Function calculated only on the set of masked view indices (M) to ensure reconstructed features are semantically meaningful and aligned with the original feature distribution, using Mean Squared Error (MSE): LMMR = 1/M Σ i∈M (ˆfi − f gti)2 / 2.

  2. The Feature Memory Module (FMM) addresses semantic ambiguity by introducing a learnable memory bank that stores class-level semantic prototypes using two strategies: Single-Proto and Multi-Proto. The Single-Proto strategy maintains one global centroid per class, updated via a momentum moving average to filter noise. The Multi-Proto strategy learns multiple sub-prototypes per class to model intra-class variance, retrieving them based on cosine similarity (sk j = x · mk j / xmk j) and weighting them using a softmax function with temperature τ.

  3. Finally, the FMM injects this retrieved semantic knowledge back into the 3D volume by adding a residual correction to the reconstructed features: x' = x + Σ k=1 (P(k) × Σ j=1 αk,j mk j), where P(k) is the predicted class probability and αk,j are the retrieval weights.

Experimental Results and Contributions

Extensive experiments on the nuScenes dataset demonstrate that M2-Occ significantly enhances robustness without compromising full-view performance. Under the safety-critical “missing back view” setting, M2-Occ improves the IoU by 4.93% compared to the baseline, effectively recovering geometry in the rear blind spot. When up to five cameras are disabled, our framework maintains an IoU of 18.36%, whereas the baseline collapses to 13.35%.

Ablation Study Insights

The ablation study shows a clear division of labor: MMR restores geometric completeness by reconstructing structural features, while FMM enhances feature representation by imposing global class-level constraints. Introducing MMR alone improves IoU to 28.19%, confirming its role in compensating for large-scale structures. However, the FMM's contribution is nuanced; the Single-Proto strategy increases IoU to 28.38%, while the Multi-Proto variant achieves a lower IoU of 27.76%. This suggests that under missing-view conditions, fine-grained prototype assignment may introduce instability, making a single global centroid more robust than multiple subcentroids when observations are heavily corrupted.

Qualitative Findings

Qualitative results in Figure 5 show that in the “missing FRONT view” scenario, the baseline method fails to detect the road and vehicles in the front blind spot, producing fragmented geometry. In contrast, M2-Occ successfully hallucinates the drivable surface and vehicle structures, demonstrating its ability to infer occluded regions from contextual cues. While fine-grained details of small objects remain challenging to fully recover, the overall perceptual quality is significantly enhanced across diverse scenarios with missing front views.

Model Efficiency

Compared with the baseline, M2-Occ increases video memory consumption by only approximately 0.15 GB (about 2.5%) while achieving significant performance gains.

Improvements for AI systems

Here are the specific improvements for existing AI systems based on the M2-Occ framework, and what these improved systems can achieve:


  1. The proposed system (M2-Occ) should be integrated into any semantic occupancy prediction pipeline that relies on multi-camera input but is currently vulnerable to sensor failures or occlusion (e.g., SurroundOcc).

  2. Implement the Multi-view Masked Reconstruction (MMR) module to recover missing view features by leveraging spatial overlaps between adjacent cameras, using a transformer-based decoder guided by learnable mask tokens and positional embeddings.

  3. Integrate the Feature Memory Module (FMM) to refine the reconstructed 3D voxel representations by retrieving global semantic prototypes (using either Single-Proto or Multi-Proto strategies) and injecting these priors as residual corrections based on predicted class probabilities.

The improved AI system can achieve the following specific capabilities:

  1. It will significantly enhance the reliability of 3D geometric and semantic understanding in autonomous driving scenarios where camera inputs are incomplete (e.g., due to temporary occlusion, physical damage, or communication dropouts).

  2. It will enable robust perception in safety-critical missing back-view situations by effectively hallucinating missing structural information (like road surfaces and vehicle volumes) based on contextual cues from neighboring cameras, leading to a measurable improvement in geometric completeness (IoU).

  3. It will maintain semantic coherence under incomplete observations by using learned global semantic prototypes to stabilize ambiguous voxel features, ensuring that reconstructed objects retain their correct class attributes even when visual evidence is partial.

  4. It will demonstrate superior robustness against catastrophic sensor failures (e.g., up to five missing views), where the system can preserve essential structural information and stabilize occupancy predictions where baseline models collapse or degrade severely.

  5. It will allow for a more reliable perception of large-scale structures (roads, drivable areas) by effectively compensating for missing views, even if fine-grained semantic details of small objects remain challenging to recover under severe dropout conditions.

Sources

Related papers