SILSA: Sliding-Window Slice Latents for Topology-Preserving High-Resolution 3D Generation
cs.CV, cs.AI, cs.LG
Submitted: 2026-10-01
Updated: 2026-10-01
Project page: https://plan-lab.github.io/silsa
Terminology
Sources
- Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation
- Demystifying MMD GANs
- V3D: Video Diffusion Models are Effective 3D Generators
- Self-supervised Learning of Hybrid Part-aware 3D Representations of 2D Gaussians and Superquadrics
- 3DGen: Triplane Latent Diffusion for Textured Mesh Generation
- LRM: Large Reconstruction Model for Single Image to 3D
- Topology-Aware Latent Diffusion for 3D Shape Generation
- SPAR3D: Stable Point-Aware Reconstruction of 3D Objects from Single Images
- STITCH: Surface reconstrucTion using Implicit neural representations with Topology Constraints and persistent Homology
- Shap-E: Generating Conditional 3D Implicit Functions
- CraftsMan3D: High-fidelity Mesh Generation with 3D Native Generation and Interactive Geometry Refiner
- Flow Matching for Generative Modeling
- ReconX: Reconstruct Any Scene from Sparse Views with Video Diffusion Model
- SyncDreamer: Generating Multiview-consistent Images from a Single-view Image
- DINOv2: Learning Robust Visual Features without Supervision
- DreamFusion: Text-to-3D using 2D Diffusion
- Zero123++: a Single Image to Consistent Multi-view Diffusion Base Model
- MVDream: Multi-view Diffusion for 3D Generation
- Efficient Betti Matching Enables Topology-Aware 3D Segmentation via Persistent Homology
- Consistent123: Improve Consistency for One Image to 3D Object Synthesis
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models