CRePE: Curved Ray Expectation Positional Encoding for Unified-Camera-Controlled Video Generation
cs.CV, cs.AI, cs.LG
Submitted: 2026-05-13
Updated: 2026-09-05
Comments: ECCV 2026 Workshop on 3D in the Era of World Models (3DWM). 29 pages, including supplementary material
License: http://creativecommons.org/licenses/by/4.0/
The gist: Video world models should predict future appearance in a way that remains consistent with 3D scene structure, camera motion, and lens geometry.
Terminology
Abstract
Video world models should predict future appearance in a way that remains consistent with 3D scene structure, camera motion, and lens geometry. Existing attention-level camera encodings, however, either describe each token only by its viewing ray---without locating scene content along that ray---or assume pinhole projection, limiting camera control under wide-angle and fisheye lenses. We introduce Curved Ray Expectation Positional Encoding (CRePE), which represents each image token as a depth-aware distribution along its Unified Camera Model (UCM) ray and integrates the expected rotary positional phasor along the curved path this distribution traces when projected into each query view. CRePE is realized through a lightweight Geometric Attention Adapter on a frozen video diffusion transformer, with pseudo radial-distance supervision from a monocular geometry foundation model serving as a stabilizing anchor rather than an inference-time input. CRePE improves camera-control, lens, and orientation fidelity across pinhole, wide-angle, and fisheye settings, and transfers zero-shot to unseen real fisheye and diverse pinhole videos. Through Radial MixForcing, the same positional pathway further accepts externally supplied radial maps, enabling scene-geometry-conditioned generation and source-video motion transfer that follow the supplied geometry more faithfully than dedicated depth-conditioned baselines. CRePE thus offers a compact interface that unifies camera control, implicit 3D scene state, and external geometry control for video world models.
Sources
- RayRoPE: Projective Ray Positional Encoding for Multi-view Attention
- Unified Camera Positional Encoding for Controlled Video Generation
- Scalable Diffusion Models with Transformers
- Wan: Open and Advanced Large-Scale Video Generative Models
- ReCamMaster: Camera-Controlled Generative Rendering from A Single Video
- Cameras as Rays: Pose Estimation via Ray Diffusion
- RoFormer: Enhanced Transformer with Rotary Position Embedding
- EscherNet: A Generative Model for Scalable View Synthesis
- GTA: A Geometry-Aware Attention Mechanism for Multi-View Transformers
- Cameras as Relative Positional Encoding
- Light Field Networks: Neural Scene Representations with Single-Evaluation Rendering
- Learning Neural Light Fields with Ray-Space Embedding Networks
- Scene Representation Transformer: Geometry-Free Novel View Synthesis Through Set-Latent Scene Representations
- CAT3D: Create Anything in 3D with Multi-View Diffusion Models
- DUSt3R: Geometric 3D Vision Made Easy
- Grounding Image Matching in 3D with MASt3R
- VGGT: Visual Geometry Grounded Transformer
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models