Omni-Diffusion-Distill: Few-Step Distillation of Unified Multimodal Diffusion Large Language Models
cs.CV, cs.LG
Submitted: 2026-10-07
Updated: 2026-10-07
Terminology
Sources
- LLaDA2.0-Uni: Unifying Multimodal Understanding and Generation with Diffusion Large Language Model
- Block Diffusion: Interpolating Between Autoregressive and Diffusion Language Models
- Uni-DAD: Unified Distillation and Adaptation of Diffusion Models for Few-step Few-shot Image Generation
- Halton Scheduler For Masked Generative Image Transformer
- A Continuous Time Framework for Discrete Denoising Models
- BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset
- RegionE: Adaptive Region-Aware Generation for Efficient Image Editing
- Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling
- Microsoft COCO Captions: Data Collection and Evaluation Server
- dParallel: Learnable Parallel Decoding for dLLMs
- Self-Forcing++: Towards Minute-Scale High-Quality Video Generation
- Emerging Properties in Unified Multimodal Pretraining
- Beyond Autoregression: Fast LLMs via Self-Distillation Through Time
- Taming Transformers for High-Resolution Image Synthesis
- MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
- GenEval: An Object-Focused Framework for Evaluating Text-to-Image Alignment
- DiffuSeq: Sequence to Sequence Text Generation with Diffusion Models
- SSD-LM: Semi-autoregressive Simplex-based Diffusion Language Model for Text Generation and Modular Control
- Distillation of Discrete Diffusion through Dimensional Correlations
- Distilling the Knowledge in a Neural Network
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models