LVSPM: Long Sequence View Synthesis and Pose Estimation Model
cs.CV
Submitted: 2026-10-07
Updated: 2026-10-07
Project page: https://burningdust21.github.io/Projects/LVSPM
Terminology
Sources
- TTT3R: 3D Reconstruction as Test-Time Training
- Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality
- InstantSplat: Sparse-view Gaussian Splatting in Seconds
- Mamba: Linear-Time Sequence Modeling with Selective State Spaces
- No Pose at All: Self-Supervised Pose-Free 3D Gaussian Splatting from Sparse Views
- LEAP: Liberate Sparse-view 3D Modeling from Camera Poses
- RayZer: A Self-supervised Large View Synthesis Model
- AnySplat: Feed-forward 3D Gaussian Splatting from Unconstrained Views
- Few-View Object Reconstruction with Unknown Categories and Camera Poses
- RelPose++: Recovering 6D Poses from Sparse-view Observations
- Depth Anything 3: Recovering the Visual Space from Any Views
- MVSGaussian: Fast Generalizable Gaussian Splatting Reconstruction from Multi-View Stereo
- R2D2: Repeatable and Reliable Detector and Descriptor
- GLU Variants Improve Transformer
- Splatt3R: Zero-shot Gaussian Splatting from Uncalibrated Image Pairs
- Learning to (Learn at Test Time): RNNs with Expressive Hidden States
- MeshLRM: Large Reconstruction Model for High-Quality Meshes
- SiNeRF: Sinusoidal Neural Radiance Fields for Joint Pose Estimation and Scene Reconstruction
- FreeSplatter: Pose-free Gaussian Splatting for Sparse-view 3D Reconstruction
- No Pose, No Problem: Surprisingly Simple 3D Gaussian Splats from Sparse Unposed Images
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models