Dynamic Image Prompt Adapter for Scalable Zero-shot Personalized Text-to-Image Generation

arXiv:2512.09814 · cs.CV · Submitted 2025-12-10 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "Dynamic Image Prompt Adapter for Scalable Zero-shot Personalized Text-to-Image Generation".

Tom: Personalized Text-to-Image (PT2I) generation aims to produce customized images based on reference images, and this work introduces DynaIP, a cutting-edge plugin designed to enhance fine-grained concept fidelity,

Jane: First, who's behind it and why it matters.

Paper summary: Tom: Building on that, Jane, can you break down what exactly the paper claims regarding the core components of DynaIP and why this setup matters so much for zero-shot personalization? I'm interested in how they structure their approach.

Jane: Okay, so they propose DynaIP as a plugin designed to enhance fine-grained concept fidelity while improving the balance between Concept Preservation and Prompt Following, plus making it scalable for multi-subject compositions within state-of-the-art Text-to-Image multimodal diffusion transformers. Essentially, the paper shows how to take a single subject personalization method and extend it directly to multi-subject personalization using a simple mask-guided injection of features.

Lu: The summary mentions that they reveal the visual encoder itself is a critical factor in fine-grained concept preservation, which is an interesting insight because it points toward where we might need to focus our architectural improvements next.

Meng: So, when they talk about extending single-subject personalization to multi-subject personalization through mask-guided region-level feature injection, that sounds like a very direct way to solve the scalability problem they identify in other adapter methods. How does that mechanism actually function operationally?

Jane: It involves computing cross attention between the noisy image tokens and each reference image token, and then adding those outputs to the corresponding regions of the text CA output features guided by their masks, using a formula like "DCA''(T, X, C) = XMMA + X N i=one λi · Mi · CA(X, Ci)". This process is what allows them to inject the specific visual information from each reference image into the generation process in a controlled way.

Tom: That sounds like a very specific and targeted injection method, Jane. It's not just slapping features on; it’s region-level guidance which suggests they are being very precise about *where* the information goes.

Lu: The paper also highlights that they leverage the hierarchical features of commonly used CLIP models to capture visual data at different levels of detail, which leads into their second key innovation: the Hierarchical Mixture-of-Experts Feature Fusion Module.

Meng: That fusion module sounds like it's designed to handle those diverse granularities effectively. Does this mean they can choose how much detail they want to preserve for a specific subject or concept?

Jane: Exactly, Meng. The HMoE-FFM uses a routing mechanism to dynamically calibrate fusion coefficients based on the characteristics of each reference image, allowing users to manually adjust the fusion coefficients of experts’ outputs, which gives them flexible control over how much fine-grained detail they want to keep in the final image.

Lalam: This level of granular control is fascinating because it means we aren't stuck with a single fidelity setting; we can tailor the visual complexity precisely to what's needed for a specific artistic vision.

Conclusion: Tom: So, wrapping up this discussion on the "Dynamic Image Prompt Adapter for Scalable Zero-shot Personalized Text-to-Image Generation," I think the title itself really tells you a lot about what they set out to accomplish—it’s all about dynamic adaptation and scaling personalization.

Jane: It really is, Tom. The authors have successfully shown how their proposed DynaIP addresses the major issues we see in current methods, specifically tackling that tough trade-off between concept preservation and prompt following while simultaneously solving the problem of extending single-subject generation to multi-subject generation without needing new training data for every combination.

Lu: The implication here is that if this works as claimed, it suggests that future large scale generative models will be far more adaptable to specific user needs and complex compositional requests than what we currently have available in the public domain.

Meng: From a practical standpoint, the ability to generate images of multiple subjects consistently with high fidelity based on just one subject’s training data is significant because it drastically reduces the need for massive, expensive datasets tailored for every new scene or character.

Lalam: I see this as a major step forward in making AI creative tools truly personal; imagine generating a whole family portrait where everyone looks exactly like them, and we only trained on one person!

Tom: That's what I’m hearing, Jane—it moves us toward a generation workflow where users can input a single reference and get highly customized results for various applications right away. The DynaIP paper shows the path to making those complex scenes accessible in real-time.

Zhizhong Wang, Tianyi Chu, Zeyi Huang, Nanyang Wang, Kehan Li

Central Media Technology Institute, Huawei

cs.CV

Submitted: 2025-12-10

Updated: 2026-09-28

Code: https://github.com/black-forest-labs/flux

Importance score: 91/100

The gist: Personalized Text-to-Image (PT2I) generation aims to produce customized images based on reference images, and this work introduces DynaIP, a cutting-edge plugin designed to enhance fine-grained

Key concepts

Dynamic Decoupling Strategy (DDS)
This strategy dynamically separates concept-specific information from concept-agnostic information within reference images. During inference, it isolates the noisy image branch to focus on capturing unique subject details while allowing the text branch to handle general elements like posture or lighting, improving balance and scalability.
Hierarchical Mixture-of-Experts Feature Fusion Module (HMoE-FFM)
This module uses layer-specific expert networks to process hierarchical features from CLIP. A routing mechanism dynamically adjusts fusion coefficients based on the reference image, allowing users to manually control how much detail is preserved for different parts of the image.
Multi-Subject Personalization
DynaIP extends personalization to generate images with multiple subjects without needing extra training data. It achieves this by injecting reference features into specific regions of the text output guided by masks, using the DDS to manage visual integration inconsistencies across different subjects.

Terminology

Summary

Personalized Text-to-Image (PT2I) generation aims to produce customized images based on reference images, and this work introduces DynaIP, a cutting-edge plugin designed to enhance fine-grained concept fidelity, balance between Concept Preservation (CP) and Prompt Following (PF), and scalability for multi-subject compositions in state-of-the-art Text-to-Image multimodal diffusion transformers.

Key Findings

(This section is derived from the abstract and conclusion)

The key finding is that MM-DiT inherently exhibit decoupling learning behavior when injecting reference image features into its dual branches via cross attentions. The noisy image branch selectively captures the concept-specific information of the reference image, while the text branch learns concept-agnostic information. Based on this, they design an innovative Dynamic Decoupling Strategy (DDS) that removes the interference of concept-agnostic information during inference, significantly enhancing CP·PF balance and bolstering scalability. Furthermore, they reveal that hierarchical features of commonly used CLIP can capture visual information at diverse granularity levels, leading to a novel Hierarchical Mixture-of-Experts Feature Fusion Module (HMoE-FFM) that elevates fine-grained concept fidelity while providing flexible control over visual granularity.

Dynamic Decoupling Strategy (DDS)

The DDS is designed to improve the CP·PF balance and subject scalability by dynamically disentangling concept-specific information from concept-agnostic information within reference images. This strategy leverages an intrinsic characteristic of MM-DiT, where the text branch interacts with the noisy image branch through cross attentions.

  1. In training, reference image features are injected into both branches simultaneously: CA([T, X], C) = sof tmax(QK⊤√C d)VC.

  2. The output is then modified to enhance concept-specific capture: DCA'(T, X, C) = [T MMA, XMMA] + λ · CA([T, X], C).

  3. During inference, the strategy involves dynamically removing concept-agnostic information by performing cross-attention exclusively with the noisy image branch: CA(X, C) = sof tmax(QXK⊤√C d)VC. This isolates the noisy image branch to focus on capturing concept-specific information of the reference image, such as the subject’s ID and unique appearance, while allowing the text branch to learn concept-agnostic information, such as posture, perspective, and illumination.

Hierarchical Mixture-of-Experts Feature Fusion Module (HMoE-FFM)

The HMoE-FFM is introduced to fully leverage the hierarchical features of CLIP by addressing the limitation where deep features incur substantial loss of detailed information. It utilizes a routing mechanism to dynamically calibrate fusion coefficients based on the characteristics of each reference image.

  1. The module deploys "layer-specific expert networks Expertl—each comprising a linear layer with GELU (Hendrycks & Gimpel, 2016) activation and layer normalization—to process the full tokens of hierarchical features."

  2. A routing module dynamically calibrates fusion coefficients: wl = Routel(ΦCLS l), l ∈ [Low, M id, High], X l wl = 1.

  3. The final output feature is a weighted fusion: ΦF used = X l wl · el, l ∈ [Low, M id, High]. This approach allows users to manually adjust the fusion coefficients of experts’ outputs, enabling flexible control over the granularity of concept preservation according to specific needs.

Multi-Subject Personalization

DynaIP extends Single-Subject PT2I (SS-PT2I) directly to Multi-Subject PT2I (MS-PT2I) without requiring additional training on multi-subject datasets. This is achieved through a straightforward mask-guided region-level feature injection.

  1. The process involves computing cross attention between noisy image tokens and each reference image token: CA(X, Ci) as specified in Eq. (2).

  2. These outputs are then added to the corresponding regions of the text CA output features, guided by their respective masks: DCA''(T, X, C) = XMMA + X N i=1 λi · Mi · CA(X, Ci), where Mi is a binary mask specifying replacement regions.

  3. The DDS helps mitigate visual integration inconsistency when composing multiple subjects by disentangling concept-agnostic information, thereby significantly boosting the subject scalability.

Experimental Validation and Control

Extensive experiments across single- and multi-subject PT2I tasks verify that DynaIP outperforms existing approaches, marking a notable advancement in the field.

  1. The method demonstrates superior performance in attaining fine-grained concept fidelity and striking an optimal balance between CP and PF compared to baselines like IP-Adapter-Plus or FLUX.

Improvements for AI systems

Here are the specific improvements that can be made to AI systems using the DynaIP framework, along with a description of what these improved systems will be able to do:


) 1. Enhanced Concept Preservation (CP) and Prompt Following (PF) in Zero-Shot PT2I:

The system will achieve significantly higher fidelity when generating images based on reference images, as it dynamically disentangles concept-specific information (subject identity, texture) from concept-agnostic information (posture, illumination). This means the generated output will more accurately reflect the intricate details of the reference subject while strictly adhering to complex textual prompts regarding style and mood.

) 2. Fine-Grained Visual Control over Concept Fidelity:

By implementing a Hierarchical Mixture-of-Experts Feature Fusion Module (HMoE-FFM), users gain explicit, manual control over visual granularity. The system will allow users to fine-tune the balance between preserving subtle textures (low/mid layers) and capturing high-level semantic attributes (deep layers), enabling precise manipulation of the image's visual complexity—from capturing fine facial details to rendering broad stylistic elements like Impressionist swirls.

) 3. Robust Scalability for Multi-Subject Personalization:

The system will overcome the limitation of requiring massive, pre-trained multi-subject datasets. Using mask-guided region-level feature injection, the model can seamlessly integrate multiple distinct subjects (e.g., a man and a character) into a single scene based solely on reference images and masks, ensuring logical spatial relationships and harmonious compositions without needing additional subject training.

) 4. Adaptable Style Transfer via Dynamic Routing:

The HMoE-FFM's dynamic routing mechanism allows the system to adapt its feature fusion strategy in real-time based on the input image's characteristics. This enables style-aware personalization, where the system automatically prioritizes low-level texture experts for brushstroke consistency (e.g., Impressionist) or high-level semantic experts for mood transfer (e.g., melancholy tones), leading to more aesthetically coherent and contextually accurate style transfers than methods relying on simple addition or concatenation.

) 5. Enhanced Editability and Controllability:

The decoupled learning behavior, combined with the dynamic decoupling strategy (DDS), ensures that prompt modifications do not cause catastrophic copy-paste artifacts. This results in a system where users can make targeted edits—such as changing only the lighting or the style of one subject—without degrading the fidelity of other elements, leading to superior controllability in complex iterative editing workflows.

) 6. Optimized Feature Selection for Subject Fidelity:

By analyzing layer-specific reconstruction capabilities (as shown in ablation studies), the system can be optimized to select optimal CLIP feature layers (e.g., layer 10 and 17) for specific tasks, ensuring that the extracted visual information is optimally balanced—capturing necessary fine-grained details without suffering from the noise or abstraction associated with overly deep or overly shallow features.

Sources

Related papers