SurGe: Improved Surface Geometry in Point Maps

arXiv:2605.31577 · cs.CV · Submitted 2026-05-29 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "SurGe: Improved Surface Geometry in Point Maps".

Tom: Recent feedforward 3D reconstruction methods predict point maps and estimate global 3D geometry remarkably well, but these predictions still exhibit inaccurate local surface geometry,

Jane: First, who's behind it and why it matters.

Paper summary: Tom: So, to recap where we are is that SurGe is tackling how recent feedforward three dee reconstruction methods predict point maps and estimate global geometry very well, but they still fail on local surface geometry, which is visible but not strongly reflected in typical metrics <ref:2605.31577#pg0,recent feedforward 3D reconstruction methods predict point maps and estimate global>.

Jane: Essentially, the authors introduce a new point map normal metric called MAEnormal to better evaluate that local surface orientation induced by neighboring three dee predictions <ref:2605.31577#pg0,local surface orientation induced by neighboring 3D predictions>.

Lu: They propose two main things to fix this: first, a point gradient matching loss that supervises depth-normalized three dee finite differences for local consistency <ref:2605.31577#pg0,a point gradient matching loss that supervises depth-normalized 3D finite differences>.

Meng: And second, they have a Neighborhood Attention Decoder, which replaces standard convolutional kernels with blocks based on Neighborhood Attention to selectively aggregate local evidence without the artifacts we often see.

Lalam: It’s a clever way to focus the model's attention precisely where it needs to be—on those subtle local connections between nearby points.

Tom: Exactly, so they're not just tweaking the overall accuracy of the point map; they are specifically targeting and improving that local surface structure, which is what makes a big difference in visual quality.

Conclusion: Jane: Looking at the title, "SurGe: Improved Surface Geometry in Point Maps," it clearly signals that this research is moving beyond just getting a generally accurate three dee model to making sure the surface details are actually correct on a local level <ref:2605.31577#pg0,SurGe: Improved Surface Geometry in Point Maps>.

Lu: The authors, Karim Knaebel and colleagues, have developed a new evaluation metric called MAEnormal which they claim reflects local surface quality much better than the standard metrics we usually rely on.

Meng: From an engineering standpoint, this means that when we deploy these reconstruction models, we can use this new metric to monitor exactly where the model is failing locally on thin structures.

Lalam: The implication for culture and AI systems is that if we can reliably measure and enforce better local surface geometry through training losses like the point gradient matching loss mentioned in the paper, we could start building AI models that inherently produce more geometrically sound visualizations for everything from architectural modeling to robotics.

Tom: It boils down to this: SurGe isn't just about getting a slightly better average error score; it's about explicitly teaching the model how to capture the orientation induced by its neighbors, leading to visibly smoother and more accurate local surfaces across eight different benchmarks.

Karim Knaebel, Gonzalo Martin Garcia, Christian Schmidt, Ilya Fradlin, Lucas Nunes, Daan de Geus

RWTH Aachen University

cs.CV

Submitted: 2026-05-29

Updated: 2026-10-02

Importance score: 89/100

The gist: Recent feedforward 3D reconstruction methods predict point maps and estimate global 3D geometry remarkably well, but these predictions still exhibit inaccurate local surface geometry, which is

Key concepts

MAEnormal
This is a new evaluation metric that directly measures the local surface orientation derived from differences between neighboring points. Unlike standard metrics that only look at average positioning errors, MAEnormal assesses how well the model captures the actual direction of the surface locally, which is crucial for thin structures.
Lpgm Loss
This is a specific loss function designed to supervise local surface structure in point maps. It compares depth-normalized 3D finite differences between neighboring points, making it 'scale-invariant' by normalizing the comparison using the depth of the nearer endpoint. This loss helps enforce consistency and suppress unwanted oscillations in the predicted local geometry.
Neighborhood Attention Decoder (NAD)
This is a specialized decoder architecture that replaces standard convolutional layers with blocks based on Neighborhood Attention. Instead of processing every pixel globally, it selectively aggregates local evidence from surrounding points, allowing the model to better capture fine details and thin structures without creating artifacts.

Terminology

Summary

Recent feedforward 3D reconstruction methods predict point maps and estimate global 3D geometry remarkably well, but these predictions still exhibit inaccurate local surface geometry, which is clearly visible qualitatively but only weakly reflected in common metrics. This paper introduces SurGe, a model that improves local point map and point map normal evaluations by incorporating a novel metric and architectural changes to better capture the orientation induced by neighboring 3D predictions.

How it works

The authors address the weakness of standard point map metrics, which only measure average point positioning error after global alignment, by introducing a point map normal metric that directly evaluates local surface orientation induced by neighboring point differences. This metric is quantified as MAEnormal, which compares the normals induced by neighboring point differences rather than just the magnitude of underlying residuals.

The proposed solution involves two complementary components:

  1. A point gradient matching loss (Lpgm) that supervises depth-normalized 3D finite differences to enforce local consistency and suppress oscillations. This loss is adapted from log-depth gradient matching, making it scale-invariant for each neighboring pair by normalizing the comparison by the nearer endpoint depth.

  2. A Neighborhood Attention Decoder (NAD) that replaces fixed convolutional kernels with blocks based on Neighborhood Attention (NA). This allows the decoder to selectively aggregate local evidence unlike convolutional decoders, without incurring the cost of full-resolution self-attention and without producing patch-aligned artifacts.

Key Contributions

The paper enumerates its main contributions as follows:

  1. An evaluation metric based on point map normals that better reflects local surface quality than standard point map metrics (MAEnormal).

  2. A scale-invariant point gradient matching loss (Lpgm), inspired by log-depth gradient matching [23], for supervising local surface structure in point maps.

  3. A decoder based on Neighborhood Attention for dense point map prediction that improves thin structures and local surface geometry (NAD, Sec. 3.1).

Model Architecture Details

SurGe combines a DINOv2 [29]-initialized ViT [9] encoder with the Neighborhood Attention Decoder (NAD). The NAD consists of five stages that progressively upsample the features, with each stage built from nl NAD blocks. Each NAD block is a Transformer-style residual block featuring a Neighborhood Attention layer and a pointwise Feed-Forward Network (FFN). The model utilizes window-matched RoPE [40] on queries and keys to provide consistent representation of relative offsets inside local windows, and it omits the usual pre-attention and pre-FFN LayerNorm layers.

Training Objectives

The full dense-label objective function is defined as:

L = Lglob + Lloc,4 + Lloc,16 + Lloc,64 + 10Lpgm.

Where:

Lglob

is the global affine-invariant point map loss with ROE [50] alignment.

(Lloc,

are local patch losses with diameters of 1/4, 1/16, and 1/64 of the image diagonal.

The authors note that they adapt supervision to label quality: for synthetic labels, they use the full objective (including Lpgm), while for SfM labels, they omit Lloc,64 and Lpgm.

Experimental Results

SurGe achieves the best average rank for global point map AbsRelglob and consistently improves local point map and point map normal evaluations across eight zero-shot monocular geometry benchmarks. Specifically, SurGe gives the lowest AbsRelloc on every evaluated dataset, and MAEnormal shows superior performance in reflecting local surface orientation compared to state-of-the-art methods like MoGe-2 and InfiniDepth [63], which produce oscillatory or bending artifacts on thin structures. The results demonstrate that SurGe improves local surface quality while preserving state-of-the-art global point map accuracy.

Ablation Study Insights

Ablations confirm the effectiveness of both components:

Decoder Design:

The NAD performs better than DPT heads, ConvStack, and ViT decoders in local evaluations (AbsRelloc and MAEnormal), whereas the ViT decoder suffers from patch-level artifacts. The ConvStack-L baseline shows moderate improvement over standard ConvStack.

Surface Loss:

The proposed Lpgm loss is shown to noticeably outperform both point map normals (Lnormal) and log-depth gradient matching (Lgm) in most cases, which is attributed to the fact that the Lpgm signal remains in the same space as the global loss.

Conclusion

SurGe presents a monocular point map model designed to improve local surface geometry rather than only average point accuracy.

Improvements for AI systems

Here are the specific improvements that can be made to existing point map prediction systems by adopting SurGe, along with what these improved AI systems will be able to achieve:


  1. The integration of the proposed Neighborhood Attention Decoder (NAD) into a feedforward 3D reconstruction pipeline.

  2. The implementation of the scale-invariant Point Gradient Matching Loss (Lpgm) as a critical component of the training objective.

  3. The adoption of the Point Map Normal Metric (MAEnormal) for explicit, fine-grained evaluation of local surface orientation, rather than relying solely on global point position errors (AbsRelglob).

These improvements will enable the following capabilities in AI systems:

  1. Predicting and reconstructing thin structures with high fidelity, such as street signs, chair legs, and faucets.

  2. Producing 3D models that exhibit significantly smoother surfaces and sharper geometric detail compared to current state-of-the-art methods (like MoGe or InfiniDepth), which often produce oscillatory or bending artifacts in local geometry.

  3. Achieving state-of-the-art performance in both global 3D accuracy and local surface quality simultaneously, ensuring that small errors do not compromise the overall scene structure.

  4. Developing more robust perception systems for autonomous driving and robotics, where accurate reconstruction of fine geometric details is crucial for collision avoidance and interaction with thin objects in complex environments.

  5. Creating a new generation of geometry evaluation metrics (MAEnormal) that explicitly quantify local surface coherence, allowing researchers to make targeted improvements on specific geometric defects in future models.

Sources

Related papers