Proxy2World: Learning to Generate Worlds From Lightweight Proxies without Seeing Them
cs.CV
Submitted: 2026-09-28
Updated: 2026-09-28
Project page: https://dumdumgura.github.io/proxy2world
Terminology
Sources
- Cosmos-Transfer1: Conditional World Generation with Adaptive Multimodal Control
- Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets
- UniVidX: A Unified Multimodal Framework for Versatile Video Generation via Diffusion Priors
- Video Models as Native 4D Renderers: World-Grounded Conditioning from Animated Mesh
- Control-A-Video: Controllable Text-to-Video Diffusion Models with Motion Prior and Reward Feedback Learning
- Code World Model: Coding Agent as World Brain
- Image Generators are Generalist Vision Learners
- Coarse-to-Real: Generative Rendering for Populated Dynamic Scenes
- CameraCtrl: Enabling Camera Control for Text-to-Video Generation
- Training-free Camera Control for Video Generation
- ViPE: Video Pose Engine for 3D Geometric Perception
- HunyuanVideo: A Systematic Framework For Large Video Generative Models
- Depth Anything 3: Recovering the Visual Space from Any Views
- OmniDirector: General Multi-Shot Camera Cloning without Cross-Paired Data
- Lyra 2.0: Explorable Generative 3D Worlds
- Gemini: A Family of Highly Capable Multimodal Models
- Wan: Open and Advanced Large-Scale Video Generative Models
- Geometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World Modeling
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models