AcFlow: Controlling Text-to-Image Diffusion Transformers via Learned Conditional Activation Flow
cs.CV, cs.AI
Submitted: 2026-09-09
Updated: 2026-09-29
Code: https://github.com/Nove1yst/AcFlow
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
The gist: Text-to-image diffusion transformers (DiTs) are powerful generators, yet direct prompting provides limited control interface for style intensity and can fail to suppress unwanted concepts.
Terminology
Abstract
Text-to-image diffusion transformers (DiTs) are powerful generators, yet direct prompting provides limited control interface for style intensity and can fail to suppress unwanted concepts. To enable these controls, we introduce AcFlow, an inference-time controller that transports intermediate layer image-token activations through a learned concept-conditioned velocity field while keeping the base DiT frozen. A textual concept description specifies the desired intervention, while the integration horizon provides a continuous control parameter. The field produces token-varying, activation-dependent updates. With parameters shared across concepts within each task family, the field supports fine-grained descriptions and generalizes to concepts unseen during training without per-concept fitting. On style control, AcFlow achieves the best style--content trade-off among the evaluated baselines in the high-style-alignment regime. At a fixed operating point, AcFlow attains style--content alignment of 0.5365/0.2860, compared with 0.4397/0.2684 for the baseline with the highest style alignment. Qualitative results demonstrate suppression of diverse concepts, including cases where direct prompting fails. Our analyses support the learned velocity field as an adaptive control mechanism, with update directions varying across tokens and depend on their activation states. Our code is available at https://github.com/Nove1yst/AcFlow.
Sources
- Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer
- Conditioned Activation Transport for T2I Safety Steering
- Dynamically Scaled Activation Steering
- MegaStyle: Constructing Diverse and Scalable Style Dataset via Consistent Text-to-Image Style Mapping
- Classifier-Free Diffusion Guidance
- LoRA: Low-Rank Adaptation of Large Language Models
- HQ-Edit: A High-Quality Dataset for Instruction-based Image Editing
- SHIFT: Steering Hidden Intermediates in Flow Transformers
- Semantic Steering for Controllable Generation: Tuning-Free Concept Erasure in Multimodal Diffusion Transformers
- Flow Matching for Generative Modeling
- DINOv2: Learning Robust Visual Features without Supervision
- Continual Diffusion: Continual Customization of Text-to-Image Diffusion with C-LoRA
- SteeringDiffusion: A Bottlenecked Activation Control Interface for Diffusion Models
- Qwen3 Technical Report
- Attention, May I Have Your Decision? Localizing Generative Choices in Diffusion Models
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models