Chameleon: Style-Content Disentangled Framework for Cross-Domain Object Compositing
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Chameleon: Style-Content Disentangled Framework for Cross-Domain Object Compositing".
Jane: Image compositing aims to seamlessly insert a foreground object into a background image, and recent advances in diffusion models have significantly enhanced the quality,
Tom: First, who's behind it and why it matters.
Title and authors: Tom: So, as we move into the first part of our discussion on "Chameleon: Style-Content Disentangled Framework for Cross-Domain Object Compositing," let's talk about what the title itself tells us about what they are trying to achieve. It really sets the stage for how complex this compositing task is.
Jane: Exactly, and when you read that title, you see they are specifically focused on disentangling style and content to solve cross-domain object compositing. They aren't just doing simple blending; they are trying to isolate the core substance of the foreground object from its visual presentation so it can be re-rendered in a new look.
Lu: The authors listed are Sukhun Ko, Soo Ye Kim, and Jihyong Oh, and seeing their affiliations with institutions like CMLab at Chung-Ang University and Adobe Research tells us this is coming from a very strong research environment with deep expertise in vision modeling.
Meng: I'm focusing on the authors because it suggests a high level of technical rigor. When you see names from established labs, it implies they’re dealing with problems that require solid mathematical foundations, which is good for robustness but maybe also slows down initial development.
Lalam: For me, the fact that they are focusing on disentanglement suggests a future where content and style can be swapped independently in AI generation. Imagine an AI that can take a photo of your dog and put it into a Van Gogh painting without altering the dog's shape—that’s what this direction is pointing toward.
The paper's summary: Tom: Okay, so we’ve established the premise; now let’s get into the actual substance of "Chameleon: Style-Content Disentangled Framework for Cross-Domain Object Compositing." They describe a novel two-stage training framework that uses a style and content disentangled encoder trained with Joint Hard Contrastive Learning to train an AI model called Chameleon.
Jane: To put that in simpler terms, the first stage is about teaching the encoder how to separate the style—things like brushstrokes or lighting—from the content—the actual shape and substance of the object. The second stage uses this disentangled knowledge inside a diffusion transformer with a Spatio-Temporal Attention Gating mechanism to inject that learned style into the background image in a spatially and temporally aware way.
Lu: The core idea of using Joint Hard Contrastive Learning to optimize both style loss and content loss simultaneously is what I find really interesting; it’s an efficient way to force the model to learn these distinct features rather than just one monolithic representation of the image.
Meng: It sounds like a lot of training overhead, but if it successfully disentangles things, it should give us much more control over the output compared to existing methods that rely on simpler blending techniques. I’m still wondering how stable those style and content tokens are during inference when dealing with very complex inputs.
Lalam: If this framework succeeds in creating a truly separable style and content representation, the cultural impact is huge because it moves us away from just "style transfer" toward true compositional manipulation, which opens up so much creative space for digital artists.
The paper's improvements: Tom: The authors highlight some specific improvements that differentiate their work, and they are pretty clear about what they’ve managed to fix compared to previous methods. They show that Chameleon preserves foreground identity and maintains consistent style, and critically, it achieves compositional plausibility through aligned geometry and grounded shadows.
Jane: That geometric alignment is huge for me because it means the inserted object doesn't just look pasted on; it actually fits into the scene in terms of perspective and lighting. And the paper specifically points out that they achieved grounded shadows, which adds a layer of realism that was missing in earlier attempts.
Lu: The results show improvements in both CLIP-I, which measures content fidelity, and CSD, style consistency. They also demonstrated that replacing the training objective with JHCL consistently improves both metrics on the AIComposer benchmark. That’s solid evidence supporting their choice of that specific learning objective for disentanglement.
Meng: Those quantitative results are what I care about most; if they're consistently better across those established metrics, it suggests a reliable method for achieving high quality in real-world scenarios, even if the training pipeline is complicated. I need to see how easily we can deploy this kind of robust performance.
Lalam: The fact that they showed improvements over sequential pipelines and commercial models in terms of both compositional plausibility and stylistic fidelity suggests that this framework offers a more holistic solution than just stitching together different tools.
Conclusion: Tom: So, to wrap up on "Chameleon: Style-Content Disentangled Framework for Cross-Domain Object Compositing," the main message is that by using a two-stage training approach with JHCL and STAG, they successfully disentangle style and content. This allows the model to generate insertions that are not only stylistically consistent but also geometrically plausible with proper shadows.
Jane: And as we conclude, this paper opens up possibilities for truly controllable cross-domain image generation where the object's identity is strongly preserved while its style is radically transformed by the new background domain. It’s a significant step toward more nuanced and realistic visual scene creation.
Lu: I think the implications are that we start moving past simple image manipulation toward systems that can understand and manipulate visual semantics at a deeper level, which could lead to entirely new creative tools in digital media.
Meng: From an engineering standpoint, the real implication is the need for more robust training pipelines like ChameleonDatasettr to ensure these results hold up across a wide variety of inputs before we can trust them for widespread use.
Lalam: I feel incredibly optimistic because this work paves a path toward AI systems that can truly understand and manipulate visual elements in a way that feels intuitive and creative, which is something we all want to see in our daily lives.
Sukhun Ko, Soo Ye Kim, Jihyong Oh
CMLab, Chung-Ang University · Adobe Research
cs.CV
Submitted: 2026-05-31
Updated: 2026-09-28
Code: https://github.com/christophschuhmann/improved-aesthetic-predictor
Project page: https://cmlab-korea.github.io/Chameleon
Importance score: 90/100
The gist: Image compositing aims to seamlessly insert a foreground object into a background image, and recent advances in diffusion models have significantly enhanced the quality, especially when the
Key concepts
- ChameleonEncoder
- This is an encoder designed to separate the style information from the content information in an image. It uses a technique called Joint Hard Contrastive Learning (JHCL) to explicitly train two different parts of the network: one for style and one for content. This disentanglement ensures that the foreground object's identity can be preserved while its appearance is adapted to a new background style.
- Joint Hard Contrastive Learning (JHCL)
- JHCL is a specific training objective used to train the ChameleonEncoder. It combines two separate contrastive learning tasks—one for style and one for content—into a single loss function. By jointly optimizing both, the model learns to isolate style features from content features more effectively, leading to better separation in the image representations.
- Spatio-Temporal Attention Gating (STAG)
- STAG is a mechanism within the diffusion transformer that controls how style is injected into both space and time during image generation. It uses learned gates based on the current timestep and spatial location to selectively apply background style tokens, ensuring that the foreground object's appearance matches its new environment consistently across the image.
- ChameleonDataset Construction
- The dataset was built using a reverse pipeline starting from real stylized composites. This method ensures that the training data is free of artifacts from generative models, as supervision is only applied to real composite images. The pipeline generates synthetic foreground images by varying camera parameters and lighting to create diverse training examples.
Terminology
Summary
Image compositing aims to seamlessly insert a foreground object into a background image, and recent advances in diffusion models have significantly enhanced the quality, especially when the foreground and background images come from different domains. The core challenge addressed by this work is cross-domain compositing, which requires preserving the identity of the foreground object while stylizing it to match a different background domain.
The gist
Chameleon is a novel two-stage training-based cross-domain compositing framework that achieves compositional plausibility and stylistic fidelity by employing a style and content disentangled encoder trained with Joint Hard Contrastive Learning (JHCL) and a Spatio-Temporal Attention Gating (STAG) mechanism within a diffusion transformer.
ChameleonDataset Construction
The paper constructs ChameleonDatasettr, the first large-scale training set for cross-domain compositing, built via a reverse data generation pipeline
that starts from real stylized composite images. This approach mitigates artifacts in stylized generations induced in prior datasets by ensuring supervision is applied only to the real-image composite (Ic), which remains free of generative artifacts by construction.
The pipeline involves five stages:
-
Object label generation using Qwen3-VL to produce object labels compatible with SAM3.
-
Object segmentation using SAM3 based on filtered labels from Qwen3-VL.
-
Data filtering via a secondary prompt to score candidates along multiple criteria and produce a binary keep/reject decision.
-
Reference generation where valid candidates are randomly cropped, padded, and passed through a reference generation model to obtain the synthetic foreground image (I′f).
-
Appearance variation generation using [40] to perturb camera parameters (azimuth, elevation) to obtain I′′f with diverse poses and lighting.
ChameleonEncoder: Style-Content Disentangled Encoder
The framework utilizes ChameleonEncoder, a style-content disentangled encoder trained with Joint Hard Contrastive Learning (JHCL).
This extends hard contrastive learning to DINOv3, explicitly disentangling its features into style and content tokens. The objective is formulated as LJHCL = LS + LC, jointly optimizing the style loss (LS) and content loss (LC).
-
The training involves two task-specific projection heads: one for style and one for content, both optimized with their own contrastive objective.
-
To achieve disentanglement, the similarity calculation is redefined using patch tokens from intermediate layers of DINOv3 (layers 18–20 for content and 12–14 for style).
-
The selection of layer ranges is justified by analyzing the
disentanglement margin
(∆) in Table 9, indicating that mid-level layers are selected for the style head Es and late layers (18–20) are selected for the content head Ec.
Cross-Domain Compositing Model with STAG
Stage 2 introduces a diffusion transformer (DiT) conditioned on the disentangled representations from ChameleonEncoder. The central mechanism is Spatio-Temporal Attention Gating (STAG), which adaptively regulates style injection in both space and time for effective cross-domain compositing.
-
STAG maps the diffusion timestep t to a sinusoidal embedding, fed into two separate MLPs to produce layer-wise gating logits for foreground (f) and background (b) regions.
-
These logits are converted to gating coefficients via the sigmoid function, and each query token is assigned a coefficient based on its spatial location using a binary foreground mask m(q).
-
The final attention at layer l is computed as A˜(l) = softmax(Q(l)K(l)⊤√dk + B(l)(t)), where the gating bias B(l)(t) is applied exclusively to keys corresponding to the background style tokens (ST (Ib)).
Evaluation and Results
The framework is evaluated on ChameleonDatasetev, a benchmark that includes diverse styles and challenging compositional scenarios,
such as stair-climbing, reflective surfaces, and lighting-aware shadows.
Quantitative results show that Chameleon outperforms state-of-the-art in-domain and cross-domain models. Specifically, the ablation study confirms that replacing the training objective with JHCL consistently improves both CLIP-I (content fidelity) and CSD (style consistency), leading to best or second-best results across all metrics
on the AIComposer benchmark. The user study on ChameleonDatasetev further validates this, showing a 71.5% win rate in Identity Preservation and a 57.7% win rate in Style Transfer Consistency compared to baselines like TF-ICON and AIComposer.
Limitations and Broader Impacts
The framework's current focus is on image-level compositing; extending it to video remains a challenge due to the continuous change in background scene, lighting, and geometry over time.
Improvements for AI systems
Based on the provided scientific paper, here are specific improvements for AI systems and what those improved systems can achieve:
) Improvements for Image Compositing AI Systems:
-
Acknowledge and Address Cross-Domain Challenges Directly:
-
Mitigate Identity Drift in Style Transfer:
-
Ensure Consistent Background Style Adaptation:
-
Achieve Geometrically Plausible Composition (Shadows and Alignment):
-
Leverage Disentangled Representations for Robust Conditioning:
) Specific Improvements and Capabilities of the Improved AI System (Chameleon Framework):
The improved system, built upon the Chameleon framework, will transition from blending-based
or frozen T2I backbone
methods to a robust, training-based generative pipeline. Its specific capabilities include:
- Acknowledge and Address Cross-Domain Challenges Directly:
The system will be explicitly trained on cross-domain data (ChameleonDatasettr) rather than relying on domain-specific blending heuristics. This allows it to handle insertions between vastly different visual domains (e.g., inserting a photorealistic object into a watercolor sketch) without requiring manual, style-specific prompt engineering for every new pair.
- Mitigate Identity Drift in Style Transfer:
By utilizing the novel two-stage training framework involving the ChameleonEncoder trained with Joint Hard Contrastive Learning (JHCL), the system explicitly disentangles style
and content
representations from a powerful encoder (DINOv3). This ensures that when transferring style, the core identity of the foreground object is preserved, preventing common issues like identity drifting,
where a prompt might change the object's fundamental structure.
- Ensure Consistent Background Style Adaptation:
The system employs Spatio-Temporal Attention Gating (STAG) within a Diffusion Transformer (DiT). This mechanism adaptively regulates how style tokens from the background are injected into the generation process, modulating this injection based on both spatial location and the diffusion timestep.
- Achieve Geometrically Plausible Composition (Shadows and Alignment):
The framework integrates positional affine transformations derived from mask information to warp foreground latent indices to their target placement, ensuring accurate geometric alignment. Furthermore, it learns to synthesize grounded shadows
by incorporating a dedicated shadow component into the training objective, leading to visually coherent lighting and texture transitions between the inserted object and the background.
- Leverage Disentangled Representations for Robust Conditioning:
The core innovation is the Style-Content Disentangled Encoder (ChameleonEncoder). Instead of relying on entangled features (like those in raw DINOv3), it uses JHCL to produce pure style and pure content tokens. These disentangled tokens are then injected into the DiT, allowing the model to selectively apply style information while maintaining fidelity to the original object's content, leading to superior stylistic fidelity compared to methods that over-stylize or ignore input masks.
) Summary of Improved System Capabilities:
The improved AI system (Chameleon) can perform high-quality, controllable cross-domain image compositing by:
-
Generating photorealistic insertions across any domain (e.g., turning a 3D render into a pixel art scene).
-
Guaranteed preservation of the foreground object's specific structure and identity, regardless of the target style.
-
Seamless integration into complex scenes by generating physically plausible lighting, shadows, and perspective alignments that match the background environment perfectly.
Sources
- IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models
- DINOv3
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- The ArtBench Dataset: Benchmarking Generative Models with Artworks
- Qwen3-VL Technical Report
- Qwen-Image Technical Report
- SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features
- Representation Learning with Contrastive Predictive Coding
- Structure-Preserving Zero-Shot Image Editing via Stage-Wise Latent Injection in Diffusion Models
- Auto-Encoding Variational Bayes
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models