AnyView: Synthesizing Any Novel View in Dynamic Scenes
cs.CV, cs.LG, cs.RO
Submitted: 2026-01-23
Updated: 2026-09-10
Comments: Project webpage: https://tri-ml.github.io/AnyView/
Code: https://github.com/nvidia-cosmos/cosmospredict2
Project page: https://tri-ml.github.io/AnyView
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Cosmos World Foundation Model Platform for Physical AI
- ReCamMaster: Camera-Controlled Generative Rendering from A Single Video
- Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets
- Dreamitate: Real-World Visuomotor Policy Learning via Video Generation
- CameraCtrl: Enabling Camera Control for Text-to-Video Generation
- View-Invariant Policy Learning via Zero-Shot Novel View Synthesis
- Wan: Open and Advanced Large-Scale Video Generative Models
- Mobi-$\pi$: Mobilizing Your Robot Learning Policy
- Depth Anything V2
- CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer
- ReCapture: Generative Video Camera Controls for User-Provided Videos using Masked Video Fine-Tuning
- CamI2V: Camera-Controlled Image-to-Video Diffusion Model
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models