SurGe: Improved Surface Geometry in Point Maps
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "SurGe: Improved Surface Geometry in Point Maps".
Tom: Recent feedforward 3D reconstruction methods predict point maps and estimate global 3D geometry remarkably well, but these predictions still exhibit inaccurate local surface geometry,
Jane: First, who's behind it and why it matters.
Paper summary: Tom: So, to recap where we are is that SurGe is tackling how recent feedforward three dee reconstruction methods predict point maps and estimate global geometry very well, but they still fail on local surface geometry, which is visible but not strongly reflected in typical metrics <ref:2605.31577#pg0,recent feedforward 3D reconstruction methods predict point maps and estimate global>.
Jane: Essentially, the authors introduce a new point map normal metric called MAEnormal to better evaluate that local surface orientation induced by neighboring three dee predictions <ref:2605.31577#pg0,local surface orientation induced by neighboring 3D predictions>.
Lu: They propose two main things to fix this: first, a point gradient matching loss that supervises depth-normalized three dee finite differences for local consistency <ref:2605.31577#pg0,a point gradient matching loss that supervises depth-normalized 3D finite differences>.
Meng: And second, they have a Neighborhood Attention Decoder, which replaces standard convolutional kernels with blocks based on Neighborhood Attention to selectively aggregate local evidence without the artifacts we often see.
Lalam: It’s a clever way to focus the model's attention precisely where it needs to be—on those subtle local connections between nearby points.
Tom: Exactly, so they're not just tweaking the overall accuracy of the point map; they are specifically targeting and improving that local surface structure, which is what makes a big difference in visual quality.
Conclusion: Jane: Looking at the title, "SurGe: Improved Surface Geometry in Point Maps," it clearly signals that this research is moving beyond just getting a generally accurate three dee model to making sure the surface details are actually correct on a local level <ref:2605.31577#pg0,SurGe: Improved Surface Geometry in Point Maps>.
Lu: The authors, Karim Knaebel and colleagues, have developed a new evaluation metric called MAEnormal which they claim reflects local surface quality much better than the standard metrics we usually rely on.
Meng: From an engineering standpoint, this means that when we deploy these reconstruction models, we can use this new metric to monitor exactly where the model is failing locally on thin structures.
Lalam: The implication for culture and AI systems is that if we can reliably measure and enforce better local surface geometry through training losses like the point gradient matching loss mentioned in the paper, we could start building AI models that inherently produce more geometrically sound visualizations for everything from architectural modeling to robotics.
Tom: It boils down to this: SurGe isn't just about getting a slightly better average error score; it's about explicitly teaching the model how to capture the orientation induced by its neighbors, leading to visibly smoother and more accurate local surfaces across eight different benchmarks.
Karim Knaebel, Gonzalo Martin Garcia, Christian Schmidt, Ilya Fradlin, Lucas Nunes, Daan de Geus
RWTH Aachen University
cs.CV
Submitted: 2026-05-29
Updated: 2026-10-02
Importance score: 89/100
The gist: Recent feedforward 3D reconstruction methods predict point maps and estimate global 3D geometry remarkably well, but these predictions still exhibit inaccurate local surface geometry, which is
Key concepts
- MAEnormal
- This is a new evaluation metric that directly measures the local surface orientation derived from differences between neighboring points. Unlike standard metrics that only look at average positioning errors, MAEnormal assesses how well the model captures the actual direction of the surface locally, which is crucial for thin structures.
- Lpgm Loss
- This is a specific loss function designed to supervise local surface structure in point maps. It compares depth-normalized 3D finite differences between neighboring points, making it 'scale-invariant' by normalizing the comparison using the depth of the nearer endpoint. This loss helps enforce consistency and suppress unwanted oscillations in the predicted local geometry.
- Neighborhood Attention Decoder (NAD)
- This is a specialized decoder architecture that replaces standard convolutional layers with blocks based on Neighborhood Attention. Instead of processing every pixel globally, it selectively aggregates local evidence from surrounding points, allowing the model to better capture fine details and thin structures without creating artifacts.
Terminology
Summary
Recent feedforward 3D reconstruction methods predict point maps and estimate global 3D geometry remarkably well, but these predictions still exhibit inaccurate local surface geometry, which is clearly visible qualitatively but only weakly reflected in common metrics. This paper introduces SurGe, a model that improves local point map and point map normal evaluations by incorporating a novel metric and architectural changes to better capture the orientation induced by neighboring 3D predictions.
How it works
The authors address the weakness of standard point map metrics, which only measure average point positioning error after global alignment, by introducing a point map normal metric
that directly evaluates local surface orientation induced by neighboring point differences. This metric is quantified as MAEnormal, which compares the normals induced by neighboring point differences rather than just the magnitude of underlying residuals.
The proposed solution involves two complementary components:
-
A point gradient matching loss (Lpgm) that supervises depth-normalized 3D finite differences to enforce local consistency and suppress oscillations. This loss is adapted from log-depth gradient matching, making it
scale-invariant for each neighboring pair
by normalizing the comparison by the nearer endpoint depth. -
A Neighborhood Attention Decoder (NAD) that replaces fixed convolutional kernels with blocks based on Neighborhood Attention (NA). This allows the decoder to
selectively aggregate local evidence unlike convolutional decoders, without incurring the cost of full-resolution self-attention and without producing patch-aligned artifacts.
Key Contributions
The paper enumerates its main contributions as follows:
-
An evaluation metric based on point map normals that better reflects local surface quality than standard point map metrics (MAEnormal).
-
A scale-invariant point gradient matching loss (Lpgm), inspired by log-depth gradient matching [23], for supervising local surface structure in point maps.
-
A decoder based on Neighborhood Attention for dense point map prediction that improves thin structures and local surface geometry (NAD, Sec. 3.1).
Model Architecture Details
SurGe combines a DINOv2 [29]-initialized ViT [9] encoder with the Neighborhood Attention Decoder (NAD). The NAD consists of five stages that progressively upsample the features, with each stage built from nl NAD blocks. Each NAD block is a Transformer-style residual block featuring a Neighborhood Attention layer and a pointwise Feed-Forward Network (FFN). The model utilizes window-matched RoPE [40] on queries and keys to provide consistent representation of relative offsets inside local windows, and it omits the usual pre-attention and pre-FFN LayerNorm layers.
Training Objectives
The full dense-label objective function is defined as:
L = Lglob + Lloc,4 + Lloc,16 + Lloc,64 + 10Lpgm.
Where:
Lglob
is the global affine-invariant point map loss with ROE [50] alignment.
(Lloc,
are local patch losses with diameters of 1/4, 1/16, and 1/64 of the image diagonal.
The authors note that they adapt supervision to label quality: for synthetic labels, they use the full objective (including Lpgm), while for SfM labels, they omit Lloc,64 and Lpgm.
Experimental Results
SurGe achieves the best average rank for global point map AbsRelglob and consistently improves local point map and point map normal evaluations across eight zero-shot monocular geometry benchmarks. Specifically, SurGe gives the lowest AbsRelloc on every evaluated dataset, and MAEnormal shows superior performance in reflecting local surface orientation compared to state-of-the-art methods like MoGe-2 and InfiniDepth [63], which produce oscillatory or bending artifacts
on thin structures. The results demonstrate that SurGe improves local surface quality while preserving state-of-the-art global point map accuracy.
Ablation Study Insights
Ablations confirm the effectiveness of both components:
Decoder Design:
The NAD performs better than DPT heads, ConvStack, and ViT decoders in local evaluations (AbsRelloc and MAEnormal), whereas the ViT decoder suffers from patch-level artifacts.
The ConvStack-L baseline shows moderate improvement over standard ConvStack.
Surface Loss:
The proposed Lpgm loss is shown to noticeably outperform both
point map normals (Lnormal) and log-depth gradient matching (Lgm) in most cases, which is attributed to the fact that the Lpgm signal remains in the same space as the global loss.
Conclusion
SurGe presents a monocular point map model designed to improve local surface geometry rather than only average point accuracy.
Improvements for AI systems
Here are the specific improvements that can be made to existing point map prediction systems by adopting SurGe, along with what these improved AI systems will be able to achieve:
-
The integration of the proposed Neighborhood Attention Decoder (NAD) into a feedforward 3D reconstruction pipeline.
-
The implementation of the scale-invariant Point Gradient Matching Loss (Lpgm) as a critical component of the training objective.
-
The adoption of the Point Map Normal Metric (MAEnormal) for explicit, fine-grained evaluation of local surface orientation, rather than relying solely on global point position errors (AbsRelglob).
These improvements will enable the following capabilities in AI systems:
-
Predicting and reconstructing thin structures with high fidelity, such as street signs, chair legs, and faucets.
-
Producing 3D models that exhibit significantly smoother surfaces and sharper geometric detail compared to current state-of-the-art methods (like MoGe or InfiniDepth), which often produce oscillatory or bending artifacts in local geometry.
-
Achieving state-of-the-art performance in both global 3D accuracy and local surface quality simultaneously, ensuring that small errors do not compromise the overall scene structure.
-
Developing more robust perception systems for autonomous driving and robotics, where accurate reconstruction of fine geometric details is crucial for collision avoidance and interaction with thin objects in complex environments.
-
Creating a new generation of geometry evaluation metrics (MAEnormal) that explicitly quantify local surface coherence, allowing researchers to make targeted improvements on specific geometric defects in future models.
Sources
- Layer Normalization
- NAF: Zero-Shot Feature Upsampling via Neighborhood Attention Filtering
- Objaverse: A Universe of Annotated 3D Objects
- Dilated Neighborhood Attention Transformer
- Generalized Neighborhood Attention: Multi-dimensional Sparse Attention at the Speed of Light
- DiNAT-IR: Exploring Dilated Neighborhood Attention for High-Quality Image Restoration
- Dilated-UNet: A Fast and Accurate Medical Image Segmentation Approach using a Dilated Transformer and U-Net Architecture
- DIODE: A Dense Indoor and Outdoor DEpth Dataset
- Pixel-Perfect Depth with Semantics-Prompted Diffusion Transformers
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models