DC-SAE: Deep Compression Semantic Autoencoder for Faster Diffusion Convergence
cs.CV
Submitted: 2026-09-30
Updated: 2026-09-30
Code: https://github.com/DAGroup-PKU/DCSAE
Project page: https://dagroup-pku.github.io/DCSAE
Terminology
Sources
- FlexTok: Resampling Images into 1D Token Sequences of Flexible Length
- Qwen3-VL Technical Report
- DC-AE 1.5: Accelerating Diffusion Model Convergence with Structured Latent Space
- Emu3.5: Native Multimodal Models are World Learners
- Unified Autoregressive Visual Generation and Understanding with Continuous Tokens
- Seed1.5-VL Technical Report
- ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment
- Ming-UniVision: Joint Image Understanding and Generation with a Unified Continuous Tokenizer
- LatentUM: Unleashing the Potential of Interleaved Cross-Modal Reasoning via a Latent-Space Unified Model
- Auto-Encoding Variational Bayes
- TC-AE: Unlocking Token Capacity for Deep Compression Autoencoders
- Autoregressive Image Generation without Vector Quantization
- DGAE: Diffusion-Guided Autoencoder for Efficient Latent Representation Learning
- Delving into Latent Spectral Biasing of Video VAEs for Superior Diffusability
- UniTok: A Unified Tokenizer for Visual Generation and Understanding
- LongCat-Next: Lexicalizing Modalities as Discrete Tokens
- Cosmos World Foundation Model Platform for Physical AI
- Latent Diffusion Model without Variational Autoencoder
- DINOv3
- UniLiP: Adapting CLIP for Unified Multimodal Understanding, Generation and Editing
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models