Camera Control for Text-to-Image Generation via Learning Viewpoint Tokens
cs.CV
Submitted: 2026-04-21
Updated: 2026-04-21
Code: https://github.com/xinsir6/ControlNetPlus
Project page: https://randdl.github.io/viewtoken_control
Terminology
Sources
- Deep Learning using Rectified Linear Units (ReLU)
- Emerging Properties in Unified Multimodal Pretraining
- SyncDreamer: Generating Multiview-consistent Images from a Single-view Image
- Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models
- T2I-Adapter: Learning Adapters to Dig out More Controllable Ability for Text-to-Image Diffusion Models
- SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis
- DreamFusion: Text-to-3D using 2D Diffusion
- Hierarchical Text-Conditional Image Generation with CLIP Latents
- Zero123++: a Single Image to Consistent Multi-view Diffusion Base Model
- MVDream: Multi-view Diffusion for 3D Generation
- MVDiffusion: Enabling Holistic Multi-view Image Generation with Correspondence-Aware Diffusion
- OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation
- Harmonizing Visual Representations for Unified Multimodal Understanding and Generation
- Show-o: One Single Transformer to Unify Multimodal Understanding and Generation
- Reconstruction Alignment Improves Unified Multimodal Models
- Qwen2 Technical Report
- TexVerse: A Universe of 3D Objects with High-Resolution Textures
- Stable Virtual Camera: Generative View Synthesis with Diffusion Models
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models