PixelUMM: Encoder-Free Unified Image and Video Understanding and Generation
cs.CV
Submitted: 2026-09-29
Updated: 2026-09-29
Code: https://github.com/black-forest-labs/flux
Project page: https://nv-tlabs.github.io/PixelUMM/Abstract
Terminology
Sources
- CM3: A Causal Masked Multimodal Model of the Internet
- Ming-Flash-Omni: A Sparse, Unified Architecture for Multimodal Perception and Generation
- Qwen2.5-VL Technical Report
- Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets
- HiDream-O1-Image: A Natively Unified Image Generative Foundation Model with Pixel-level Unified Transformer
- HunyuanImage 3.0 Technical Report
- AlignTok: Aligning Visual Foundation Encoders to Tokenizers for Diffusion Models
- BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset
- PixelFlow: Pixel-Space Generative Models with Flow
- A Simple Framework for Contrastive Learning of Visual Representations
- VUGEN: Visual Understanding priors for GENeration
- Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling
- DiP: Taming Diffusion Models in Pixel Space
- VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
- ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation
- Emu3.5: Native Multimodal Models are World Learners
- FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
- Emerging Properties in Unified Multimodal Pretraining
- Unveiling Encoder-Free Vision-Language Models
- From Pixels to Words -- Towards Native Vision-Language Primitives at Scale
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models