VGGT-Diff: Visual Geometry Meets Diffusion for Sparse-View Novel View Synthesis
cs.CV, cs.AI
Submitted: 2026-09-27
Updated: 2026-09-27
Code: https://github.com/chenkangjie1123/VGGT-Diff
Terminology
Sources
- Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets
- DreamSim: Learning New Dimensions of Human Visual Similarity using Synthetic Data
- LVSM: A Large View Synthesis Model with Minimal 3D Inductive Bias
- GeoNVS: Geometry Grounded Video Diffusion for Novel View Synthesis
- Zero123++: a Single Image to Consistent Multi-view Diffusion Base Model
- AnyView: Synthesizing Any Novel View in Dynamic Scenes
- Wan: Open and Advanced Large-Scale Video Generative Models
- VGGT-$\Omega$
- Is Attention All That NeRF Needs?
- Novel View Synthesis with Diffusion Models
- Novel View Synthesis as Video Completion
- From Rays to Projections: Better Inputs for Feed-Forward View Synthesis
- ViewCrafter: Taming Video Diffusion Models for High-fidelity Novel View Synthesis
- UniWorld-View: Large-Baseline View Synthesis via Video Diffusion Models
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models