PXDepth: Pixel-Space Modeling for Structure Preserving Monocular Depth Estimation
cs.CV, cs.AI, cs.GR
Submitted: 2026-08-17
Updated: 2026-08-31
Comments: Project Page: https://yuanzhy29.github.io/PXDepth-Page/
Project page: https://yuanzhy29.github.io/PXDepth-Page
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- ZoeDepth: Zero-shot Transfer by Combining Relative and Metric Depth
- Modeling Depth Ambiguity: A Mixture-Density Representation for Flying-Point-Free Depth Estimation
- Depth Pro: Sharp Monocular Metric Depth in Less Than a Second
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Lotus: Diffusion-based Visual Foundation Model for High-quality Dense Prediction
- Lotus-2: Advancing Geometric Dense Prediction with Powerful Image Generative Model
- On the Importance of Accurate Geometry Data for Dense 3D Vision Tasks
- Auto-Encoding Variational Bayes
- SurGe: Improved Surface Geometry in Point Maps
- MoGe-3: Fine-Detail Monocular Geometry Estimation with Self-Guided Sparse Volumetric Refinement
- From Big to Small: Multi-Scale Local Planar Guidance for Monocular Depth Estimation
- Depth Anything 3: Recovering the Visual Space from Any Views
- TartanGround: A Large-Scale Dataset for Ground Robot Perception and Navigation
- DINOv3
- Masked Depth Modeling for Spatial Perception
- DIODE: A Dense Indoor and Outdoor DEpth Dataset
- Flow-Motion and Depth Network for Monocular Stereo and Beyond
- MoGe-2: Accurate Monocular Geometry with Metric Scale and Sharp Details
- Synscapes: A Photorealistic Synthetic Dataset for Street Scene Parsing
- Pixel-Perfect Depth with Semantics-Prompted Diffusion Transformers
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models