FuseReg: Regularizing Layer Fusion Mitigates the Reconstruction-Generation Gap in Representation Autoencoders
cs.CV
Submitted: 2026-09-25
Updated: 2026-09-28
Terminology
Sources
- Laminating Representation Autoencoders for Efficient Diffusion
- Hyperspherical Autoencoder for High-Fidelity Image Reconstruction and Generation
- AlignTok: Aligning Visual Foundation Encoders to Tokenizers for Diffusion Models
- Attentive multilayer fusion for vision transformers
- The Prism Hypothesis: Harmonizing Semantic and Pixel Representations via Unified Autoencoding
- RPiAE: A Representation-Pivoted Autoencoder Enhancing Both Image Generation and Editing
- V-RAE: Rethinking Video Latent Spaces for Generation
- GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium
- REPA-E: Unlocking VAE for End-to-End Tuning with Latent Diffusion Transformers
- Flow Matching for Generative Modeling
- Improving Reconstruction of Representation Autoencoder
- Representation Alignment for Just Image Transformers is not Easier than You Think
- What matters for Representation Alignment: Global Information or Spatial Structure?
- Improved Baselines with Representation Autoencoders
- SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features
- VideoRAE: Taming Video Foundation Models for Generative Modeling via Representation Autoencoders
- Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think
- The Unreasonable Effectiveness of Deep Features as a Perceptual Metric
- Both Semantics and Reconstruction Matter: Making Representation Encoders Ready for Text-to-Image Generation and Editing
- Diffusion Transformers with Representation Autoencoders
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models