FlatLands: Generative Floormap Completion From a Single Egocentric View
cs.CV, cs.AI, cs.RO, eess.IV
Submitted: 2026-03-16
Updated: 2026-09-02
Comments: Under Review
Code: https://github.com/1ssb/Flat_Lands
License: http://creativecommons.org/licenses/by/4.0/
The gist: A single egocentric image typically captures only a small portion of the floor, yet a complete metric traversability map of the surroundings would better serve applications such as indoor navigation.
Terminology
Abstract
A single egocentric image typically captures only a small portion of the floor, yet a complete metric traversability map of the surroundings would better serve applications such as indoor navigation. We introduce FlatLands, a dataset and benchmark for single-view bird's-eye view (BEV) floor completion. The dataset contains 270,575 observations from 17,656 real metric indoor scenes drawn from six existing datasets, with aligned observation, visibility, validity, and ground-truth BEV maps, and the benchmark includes both in- and out-of-distribution evaluation protocols. We compare training-free approaches, deterministic models, ensembles, and stochastic generative models. Finally, we instantiate the task as an end-to-end monocular RGB-to-floormaps pipeline. FlatLands provides a rigorous testbed for uncertainty-aware indoor mapping and generative completion for embodied navigation.
Sources
- Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation
- MatterDoor: Sampling Zero-shot Spatio-semantic Priors using Generative Models
- FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space
- Training Deep Nets with Sublinear Memory Cost
- Flow Matching Guide and Code
- Hierarchical Text-Conditional Image Generation with CLIP Latents
- Predict-Project-Renoise: Sampling Diffusion Models under Hard Constraints
- SC-Explorer: Incremental 3D Scene Completion for Safe and Efficient Exploration Mapping and Planning
- CHOrD: Generation of Collision-Free, House-Scale, and Organized Digital Twins for 3D Indoor Scenes with Controllable Floor Plans and Optimal Layouts
- mHC: Manifold-Constrained Hyper-Connections
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models