Depth Anything in 360: Towards Scale Invariance in the Wild

arXiv:2512.22819 · cs.CV · Submitted 2025-12-28 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Depth Anything in 360: Towards Scale Invariance in the Wild".

Jane: This paper presents DA360, a panoramic-adapted version of Depth Anything V2, designed to achieve scale-invariant depth estimation from 360° panoramic images in unconstrained environments.

Tom: First, who's behind it and why it matters.

Title and authors: Tom: So we're talking about "Depth Anything in three hundred sixty: Towards Scale Invariance in the Wild." It sounds like this paper is tackling a really specific problem where standard depth estimation models just don't work well with panoramic images outside of controlled settings.

Jane: Exactly, Tom. The title tells us that the authors are focused on achieving scale invariance, which is super important because when you move from a flat perspective image to a three hundred sixty-degree panoramic view, those models struggle with knowing the actual size and position correctly.

Lu: From a theoretical standpoint, this paper is looking at closing a gap between perspective image models and what's possible with panoramic data, which is crucial for robotics and AR/VR applications that need reliable depth information everywhere.

Meng: So, to put it simply for our listeners, this research focuses on making AI systems that can accurately measure distance from three hundred sixty-degree pictures without needing tons of specific training data for every single location.

Lalam: I think the main implication here is that we can finally get depth estimation out into the wild, not just inside a lab. This work opens up possibilities for creating truly immersive AR experiences where virtual objects maintain their correct scale regardless of how you look at them.

The paper's summary: Tom: The core of the paper, "Depth Anything in three hundred sixty: Towards Scale Invariance in the Wild," is that they take a model like Depth Anything V2 and adapt it for panoramic images by learning an extra parameter to fix those scale and shift issues.

Jane: It sounds like they're taking something that works okay for perspective images, but then they're introducing a mechanism—a learnable shift parameter from the ViT backbone—to correct the output so it becomes scale-invariant.

Lu: That mechanism is key because, as the paper explains, existing methods produce depth maps that are just plausible but hard to use for building actual three dee point clouds without extra steps.

Meng: So, instead of just giving a map that might be slightly off in size or position, this new approach aims to directly give us a well-formed three dee point cloud right away.

Lalam: That direct generation capability is what really excites me for the future; it means we skip that whole complicated post-processing step and get usable three dee data instantly.

The paper's improvements: Tom: The authors introduced two main technical improvements to solve the problems they identified. First, they have this shift learning module to adjust the scale and shift.

Jane: And second, they tackled boundary artifacts by integrating circular padding into the decoder head instead of just using standard zero padding. That should keep things looking coherent around those edges of a panorama.

Lu: The authors found that this shift learning mechanism is crucial; removing it significantly degrades accuracy, especially when testing on outdoor scenarios, which is what they call in-the-wild images.

Meng: From an engineering standpoint, dealing with those boundaries is always tricky because the equirectangular projection causes inconsistencies at the seams of the panorama. Circular padding seems like a smart way to maintain spatial continuity across those boundaries.

Lalam: I see how that addresses a major weakness; it ensures that the depth map doesn't have those weird seams, which means we get much cleaner, more consistent three dee structure overall.

Conclusion: Tom: So, to wrap up the discussion on "Depth Anything in three hundred sixty: Towards Scale Invariance in the Wild," this paper shows a way to make panoramic depth estimation significantly better for real-world use by focusing on scale invariance.

Jane: They manage to reduce that uncertainty by one degree through their shift learning and padding, which leads to a substantial relative error reduction on both indoor and outdoor benchmarks.

Lu: The implication is that we can finally move beyond just indoor datasets like Matterportthree dee and Stanford2Dthree dee because they've shown promising zero-shot generalization to open-world environments, even on the newly curated Metropolis dataset.

Meng: Practically speaking, this means autonomous systems like drones or self-driving cars could get much better environmental maps from panoramic cameras without needing massive amounts of specific outdoor training data.

Lalam: For me, the biggest impact is that this method gives us a robust tool for immersive AR and VR applications; it allows for creating virtual worlds that maintain correct scale no matter the viewpoint.

Tom: It’s a solid piece of research that establishes a new state-of-the-art performance level for zero-shot panoramic depth estimation, and I think we should be really looking forward to seeing how this technology gets deployed in practical systems.

Jane: I agree, it seems like they've laid some very important groundwork for making these complex three dee reconstructions more reliable across different environments.

Lu: We just have to keep pushing the boundaries on these architectural adaptations to make them even more flexible for different types of panoramic data.

Meng: I'm curious if the computational time remains competitive, since we need fast inference for real-time applications.

Lalam: It seems competitive enough right now, and with this level of accuracy, it’s a really promising direction for how AI can interact with the physical world.

Hualie Jiang, Ziyang Song, Zhiqiang Lou, Rui Xu

Insta360 Research

cs.CV

Submitted: 2025-12-28

Updated: 2026-09-30

Comments: https://antigravity-tech.github.io/DA360

Project page: https://insta360-research-team.github.io/DA360

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 86/100

The gist: This paper presents DA360, a panoramic-adapted version of Depth Anything V2, designed to achieve scale-invariant depth estimation from 360° panoramic images in unconstrained environments.

Key concepts

Scale-Invariant Depth Estimation
This technique aims to estimate depth accurately regardless of the object's distance from the camera. It addresses a major weakness in panoramic depth models by ensuring that the estimated scale and shift parameters are learned, allowing for direct conversion into 3D structures without needing external correction steps.
Learnable Shift Module
This component uses a three-layer MLP to learn a specific shift parameter directly from the Vision Transformer backbone. This learned parameter transforms the model's standard disparity output into an estimate that is invariant to scale and shift, which is vital for producing accurate 3D data.
Circular Padding Integration
Instead of standard padding, this method uses circular padding in the decoder head. This modification ensures that depth maps maintain spatial continuity across the spherical image boundaries and effectively removes artifacts or seams that typically occur at these edges.

Terminology

Summary

This paper presents DA360, a panoramic-adapted version of Depth Anything V2, designed to achieve scale-invariant depth estimation from 360° panoramic images in unconstrained environments. It addresses the critical gap between perspective image models and panoramic depth estimation by learning a shift parameter from the ViT backbone and integrating circular padding to eliminate boundary artifacts, ultimately yielding well-formed 3D point clouds.

Motivation and Problem Statement

Panoramic depth estimation is crucial for robotics and AR/VR but lags behind perspective image methods in zero-shot generalization due to data scarcity. Existing panoramic methods often produce plausible depth maps that are difficult to convert into well-structured 3D point clouds because they remain affine-invariant, lacking the scale and shift invariance needed for direct 3D reconstruction. The paper notes that while panoramic images offer a complete field of view, current methods fail to reduce depth uncertainty by one degree, which is theoretically possible through scale-invariant estimation. Furthermore, existing evaluations are often confined to indoor datasets like Matterport3D and Stanford2D3D, leaving a lack of robust outdoor benchmarks.

Key Innovations in DA360

The core innovation of DA360 involves two primary technical components:

  1. A learnable shift module: learning a shift parameter from the ViT backbone, transforming the model’s original scale- and shift-invariant disparity output into a scale-invariant estimate. This mechanism is crucial for direct generation of well-formed 3D point clouds without cumbersome post-hoc correction.

  2. Circular padding integration: The authors address boundary issues by replacing standard zero padding with circular padding in the DPT decoder head. This ensures spatially coherent depth maps that respect spherical continuity and effectively eliminates seam artifacts caused by boundary effects and ensuring the generated depth map adheres to spherical continuity.

Methodology and Training Strategy

The framework is initialized using a strong zero-shot pinhole depth estimation model, specifically Depth Anything V2 (DAV2). The training process involves several specific steps:

  1. Shift Learning: A three-layer MLP is applied to the class token from the ViT backbone to regress the shift parameter. Ablation studies confirm that this shift learning mechanism is crucial for recovering scale-invariant disparity—its removal leads to significant accuracy degradation, particularly in outdoor scenarios.

  2. Circular Padding: The DPT decoder is modified with circular padding, which is described as being effective in maintaining spatial continuity across ERP boundaries and resolving inconsistencies at the equirectangular projection (ERP) boundaries.

  3. Scale-Invariant Loss: A novel loss function, inspired by MiDaS [33], is designed to enforce scale-invariant predictions. This involves normalizing both the predicted disparity map and ground truth using their respective robust scale estimates: dˆ pd = dpd / s(dpd), dˆ gt = dgt / s(dgt). The final loss computed is Lsi = 1/N Σ dˆ(i) pd - dˆ(i) gt.

Evaluation and Results

DA360 was evaluated on standard indoor benchmarks (Matterport3D, Stanford2D3D) and a newly curated outdoor dataset named Metropolis, derived from the Mapillary Metropolis Dataset [27]. The results demonstrate substantial gains:

**: DA360 shows over 50% and 10% relative depth error reduction on indoor and outdoor benchmarks, respectively. It also achieves about 30% relative error improvement compared to PanDA across all three test datasets. The framework establishes a new state-of-the-art performance for zero-shot panoramic depth estimation. Ablation studies confirm that learning scale-invariant depth does not underperform learning affine-invariant depth, while the former reduces one degree of uncertainty and yields directly usable 3D point clouds. Furthermore, the circular padding successfully eliminates structural inconsistencies at ERP boundaries. The paper concludes that DA360 produces more accurate depth magnitudes overall while also recovering finer local details, demonstrating superior performance over competitors like DreamCube and UniK3D in qualitative comparisons. The computational time for DA360 is 0.26s, which is competitive with other end-to-end models. The framework's strong generalization capability shows "substantial promise for enhancing real-world applications in AR/VR and autonomous systems.

Improvements for AI systems

Here are the specific improvements to AI systems derived from the DA360 framework, along with what those improved systems can achieve:


The core contribution of DA360 is transforming an existing, affine-invariant disparity estimation model (like Depth Anything V2) into a scale-invariant system capable of producing metric 3D point clouds directly, while ensuring spatial coherence across panoramic boundaries.

Here are the specific improvements and capabilities for the resulting AI systems:

The improved AI system can perform high-fidelity, zero-shot depth estimation on any 360° panoramic image (ERP format).

  1. This capability is achieved by learning a dynamic shift parameter from the Vision Transformer (ViT) backbone, which corrects the inherent scale and shift ambiguities of pinhole models.

  2. The system can output directly usable, well-structured 3D point clouds without requiring complex post-hoc correction or affine alignment steps.

  3. The system ensures spatial continuity across equirectangular projection boundaries by integrating circular padding into the decoder, eliminating seam artifacts and producing spatially coherent depth maps where traditional methods fail.

The specific capabilities gained by implementing DA360 are:

Autonomous Robotic Navigation in Open-World Environments: The ability to generate accurate, metric 3D point clouds from panoramic views allows robots (e.g., autonomous vehicles, drones) to build detailed environmental maps for global path planning and obstacle avoidance, even in novel outdoor settings where training data is scarce.

Immersive AR/VR Content Generation: Real-time rendering of geometrically accurate 3D environments directly from panoramic inputs enables high-quality virtual worlds and augmented reality experiences that maintain correct scale regardless of the camera's exact viewpoint or field of view ambiguity (FoV).

Robust Scene Understanding in Unstructured Data: By achieving zero-shot generalization to outdoor, open-world scenes (validated on the Metropolis dataset), the AI system can be deployed effectively in environments where indoor data is insufficient, overcoming the limitations of models trained only on constrained indoor benchmarks.

Enhanced Depth Accuracy for Extreme Distances: The scale-invariant loss function ensures that depth recovery remains accurate across vast ranges, preventing the numerical difficulties and coarse structure recovery observed in methods that attempt absolute depth prediction or rely solely on affine invariance.

Abstract

Panoramic depth estimation captures the complete 360 scene geometry, being essential for robotics and AR/VR applications. While perspective depth models have achieved remarkable zero-shot generalization via large-scale training, panoramic methods lag behind, especially for open-world scenes, due to data scarcity. To bridge this gap, we introduce DA360, a panoramic-adapted version of Depth Anything V2. Our key insight is that the base DAV2 model, trained on perspective images to predict affine-invariant disparity, already exhibits good zero-shot performance on panoramas. Building on this, we design a lightweight adaptation framework that (i) learns a per-image shift from the ViT class token with scale-invariant supervision, transforming affine-invariant disparity into scale-invariant disparity that directly yields well-formed 3D point clouds, and (ii) integrates circular padding into the DPT decoder to eliminate seam artifacts, ensuring spatial coherence. Fine-tuned on a combination of synthetic indoor and outdoor panoramic data, DA360 is evaluated on standard real-world indoor benchmarks and our newly curated outdoor dataset, Metropolis. Results show that DA360 not only outperforms the original DAV2 by over 50% and 12% relative error reduction indoors and outdoors, but also surpasses prior specialized methods like PanDA by about 25--35% across all tests, establishing state-of-the-art zero-shot panoramic depth estimation.

Sources

Related papers