HiLo-Token: Input-Adaptive High-Low Frequency Token Compression for Efficient Image Editing
cs.CV, cs.AI
Submitted: 2026-06-11
Updated: 2026-09-18
Comments: 16 pages, 12 figures, Patent filled
Code: https://github.com/facebookresearch/xformers
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: Creative image editing tools, such as Photoshop's Remove or Generative Fill buttons, are central to everyday customer use and account for a major share of traffic in Photoshop and Lightroom.
Terminology
Abstract
Creative image editing tools, such as Photoshop's Remove or Generative Fill buttons, are central to everyday customer use and account for a major share of traffic in Photoshop and Lightroom. However, current generative AI models face significant latency challenges, which become even more pronounced when transitioning from convolution-based U-Nets to Diffusion Transformers (DiTs). In our evaluation on hundreds of representative image editing samples spanning a wide range of mask ratios, the DiT module alone accounts for an average of 73% of the total model latency, even after being distilled from 50 timesteps down to 8 timesteps. To tackle this challenge, we propose HiLo-Token, an input-adaptive token compression framework that allocates more token budget to high-frequency, rich-context regions while assigning fewer tokens to low-frequency areas. Specifically, for the editing region specified by the user mask, we retain all tokens within a dilated mask to preserve strong locality and contextual relevance. Outside the editing region, we introduce a simple yet effective high-frequency token selection strategy based on spatial frequency to capture important local details, while using tokens from a 16x downsampled image to represent low-frequency components and preserve the blurry but global structure. Extensive experiments on production-level evaluation data validate the effectiveness of the proposed method, achieving 3.13x, 2.59x, and 1.67x DiT speedups on A100-80GB for image editing tasks across small, medium, and large mask ratio categories with average ratios of 6.38%, 15.92%, and 35.36%, respectively, without any regression in generation quality.
Sources
- Deep Compression Autoencoder for Efficient High-Resolution Diffusion Models
- Seedream 4.0: Toward Next-generation Multimodal Image Generation
- FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
- Multi-Scale Dense Networks for Resource Efficient Image Classification
- TQ-DiT: Efficient Time-Aware Quantization for Diffusion Transformers
- BK-SDM: A Lightweight, Fast, and Cheap Version of Stable Diffusion
- FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space
- Learning-to-Cache: Accelerating Diffusion Transformer via Layer Caching
- SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis
- Mixture-of-Depths: Dynamically allocating compute in transformer-based language models
- Progressive Distillation for Fast Sampling of Diffusion Models
- ToDo: Token Downsampling for Efficient Generation of High-Resolution Images
- SeedEdit 3.0: Fast and High-Quality Generative Image Editing
- ImgEdit: A Unified Image Editing Dataset and Benchmark
- PixPerfect: Seamless Latent Diffusion Local Editing with Discriminative Pixel-Space Refinement
- UniSER: A Foundation Model for Unified Soft Effects Removal
- Dynamic Diffusion Transformer
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models