GeoShrink: Accelerating Diffusion Transformers with Two Lines of Code
cs.CV
Submitted: 2026-09-27
Updated: 2026-09-27
Terminology
Sources
- MusicLM: Generating Music From Text
- All are Worth Words: A ViT Backbone for Diffusion Models
- LESA: Learnable Stage-Aware Predictors for Diffusion Model Acceleration
- $\Delta$-DiT: A Training-Free Acceleration Method Tailored for Diffusion Transformers
- POLARIS: Projection-Orthogonal Least Squares for Robust and Adaptive Inversion in Diffusion Models
- MemoGen: Can Past Experience Improve Future Text-to-Image Generation?
- A Point Set Generation Network for 3D Object Reconstruction from a Single Image
- ACE-Step: A Step Towards Music Generation Foundation Model
- T$^3$Bench: Benchmarking Current Progress in Text-to-3D Generation
- Denoising Diffusion Probabilistic Models
- HarmoniCa: Harmonizing Training and Inference for Better Feature Caching in Diffusion Transformer Acceleration
- VBench: Comprehensive Benchmark Suite for Video Generative Models
- TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization
- Adaptive Caching for Faster Video Generation with Diffusion Transformers
- Elucidating the Design Space of Diffusion-Based Generative Models
- Pick-a-Pic: An Open Dataset of User Preferences for Text-to-Image Generation
- HunyuanVideo: A Systematic Framework For Large Video Generative Models
- $Z^2$-Sampling: Zero-Cost Zigzag Trajectories for Semantic Alignment in Diffusion Models
- Delta Score Matters! Spatial Adaptive Multi Guidance in Diffusion Models
- Faster Diffusion: Rethinking the Role of the Encoder for Diffusion Model Inference
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models