EditCLIP: Representation Learning for Image Editing
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "EditCLIP: Representation Learning for Image Editing".
Tom: The paper introduces EditCLIP,
Jane: First, who's behind it and why it matters.
Paper summary: Tom: So we're talking about EditCLIP today. It’s a paper that introduces a new way of learning representations for image editing, focusing on how you can capture the transformation between an input image and its edited version in one unified space.
Jane: That sounds like it tackles the problem of using text instructions to guide complex edits in AI models, which is something we see happening a lot right now.
Tom: Exactly. The core idea behind EditCLIP is that instead of relying only on text, we learn to encode the edit itself within an embedding space by jointly looking at both the input image and the edited counterpart.
Lu: It's interesting because it frames this problem very similarly to how CLIP learns representations for images and text, aiming to learn the semantics of edits within that same CLIP space.
Meng: So, when you look at what they're trying to achieve, they want a general representation of edits so that transformations can be implicitly encoded in an embedding space.
Jane: They formulate this as transforming an input image Ii into an edited image Ie based on a textual instruction T, which is basically setting up the problem as if you were learning how to map images to text descriptions, but for edits.
Tom: Right, and the pre-training involves modifying a standard CLIP visual encoder to take a composite input where both the input and edited images are concatenated along the channel dimension.
Lalam: So they create an edit embedding E by processing that concatenation, which is then trained contrastively against the textual instruction T to align the learned editing space with the textual space.
Meng: From an engineering standpoint, training only the visual encoder while keeping the text encoder frozen means they are trying to find a structure in the visual data that relates directly to how edits are described.
Tom: Then, for actual image editing in things like diffusion models such as InstructPix2Pix, EditCLIP embeddings replace the textual instructions.
Jane: That means instead of feeding a text prompt into a model, you feed it this learned edit embedding E to guide the transformation of a new query image Iq.
Tom: During inference for a new query image Iq, the model is conditioned on that EditCLIP embedding produced from an exemplar pair, Ii and Ie, which effectively modifies how the output image Io is generated.
Lu: The evaluation part is also important; they propose two metrics: EditCLIP-to-Text similarity, EC2T, and for exemplar-based metrics, EditCLIP-to-EditCLIP, EC2EC.
Meng: These metrics are designed to quantify how well the input image transforms into the edited image while making sure those changes align with the original editing instructions.
Tom: And what they claim is that these proposed metrics align more closely with human judgments than existing CLIP-based metrics, achieving a higher correlation for both text-based and exemplar-based editing tasks.
Jane: That's significant because it suggests EditCLIP embeddings are a more reliable way to measure edit quality and structural preservation than the methods we currently use.
Tom: The quantitative results show that EditCLIP achieves state-of-the-art exemplar-based image editing with no computational overhead, which is a big deal for efficiency.
Lu: The EC2EC metric specifically performs the best among exemplar-based approaches, and they confirmed this in a user study where its winning rate was larger than fifty percent against all baselines.
Meng: For practical application, this means we have a reliable tool to automatically assess whether an AI has successfully made an edit that looks right to a human.
Jane: But the paper does point out some limitations because the current training is solely on the InstructPix2Pix dataset, which doesn't include edits like removal or deformation.
Tom: That means expanding the training data could improve both the quality and the diversity of this embedding space for different kinds of edits.
Lu: Future work they suggest involves applying EditCLIP to things like instruction caption generation or query-based editing pair retrieval, and even extensions to video and three-dimensional editing.
Jane: They also mentioned exploring more advanced training strategies, like using refined loss functions or incorporating masks as an extra channel to give better control over specific edit regions.
Tom: So, in conclusion, EditCLIP offers a representation-learning approach that captures how images transform during edits and it serves as a reliable metric for evaluating the quality and faithfulness of those edits by aligning closely with human judgment.
Jane: It seems like this method could really help speed up the development of image editing approaches because it provides an evaluation metric that matches human judgment better than what we currently have.
Conclusion: Tom: So we're wrapping up EditCLIP today, looking at how this whole thing fits together as a representation learning method for image editing.
Jane: It’s about taking an input image and its edited version, and learning a single way to describe that change using embeddings from the CLIP space.
Lu: They’re really focusing on capturing that implicit transformation, so you don't need a long text description to guide the edit anymore.
Meng: From an engineering side, they’ve built this by concatenating the input and edited images into one big picture for their visual encoder to look at.
Lalam: And we’re using that combined embedding to guide diffusion models, so instead of text instructions telling the model what to do, we use this learned edit representation.
Tom: It boils down to EditCLIP providing a unified way to encode image edits, which is pretty neat because it bypasses needing detailed natural language prompts for complex transformations.
Jane: If you’re listening and you’ve ever tried prompting an AI with a long description of a photo change, this paper shows how they can learn that transformation directly from pairs of images.
Lu: The authors are showing that their metrics for measuring these edits, like EC2EC, actually line up much better with what humans judge as good or faithful edits.
Meng: That’s the practical part I care about—if the evaluation metric is reliable, it means we can actually build systems that get better faster.
Lalam: It could mean that for image generation tasks, assessing quality becomes much more objective because this embedding space reflects the actual visual change in a way humans understand.
Tom: So, EditCLIP isn't just another model; it’s a new way to learn the semantics of editing itself within the existing powerful CLIP framework.
Jane: It moves us closer to having AI that understands *how* things look and how they can be modified visually, rather than just following simple text commands.
Lu: This opens up so many doors for future work, especially if we start thinking about applying this representation learning to video or even three dee editing later on.
Qian Wang, Aleksandar Cvejic, Abdelrahman Eldesokey, Peter Wonka
KAUST
cs.CV
Submitted: 2025-03-26
Updated: 2025-03-26
Comments: Project page: https://qianwangx.github.io/EditCLIP/
Journal ref: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 15960-15970, 2025
DOI: 10.1109/ICCV51701.2025.01481
Code: https://github.com/QianWangX/EditCLIP
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 76/100
The gist: The paper introduces EditCLIP, a novel representation-learning approach designed to capture an implicit representation of image edits by jointly encoding an input image and its edited counterpart
Key concepts
- EditCLIP
- A novel representation learning approach that creates a single embedding space where an input image and its edited version are encoded together. This allows the model to implicitly understand the transformation applied to an image based on a textual instruction.
- Exemplar-Based Editing
- A method where image editing is performed by providing a pair of images—an original and a desired edit—instead of relying solely on detailed text prompts. EditCLIP uses the joint embedding of these two images to guide the editing process in diffusion models.
- EC2T Metric
- The EditCLIP-to-Text similarity metric quantifies how well an input image transforms into its edited counterpart according to a given instruction. It measures the alignment between the combined visual embedding and the textual embedding of the instruction, serving as a reliable evaluation tool.
Terminology
Summary
The paper introduces EditCLIP, a novel representation-learning approach designed to capture an implicit representation of image edits by jointly encoding an input image and its edited counterpart within a unified embedding space, addressing limitations in text-based editing instructions and existing evaluation metrics. This method is significant because it enables exemplar-based image editing without relying on natural language descriptions for complex transformations and provides a reliable, scalable metric for assessing edit quality and structural preservation.
The gist: EditCLIP provides a unified representation of image edits by encoding the transformation between an image and its edited counterpart within the CLIP space.<ref:2503.20318#pg2>
Representation Learning for Image Editing
The core objective is to learn a general representation of edits where transformations can be implicitly encoded within an embedding space, formulated as transforming an input image Ii into an edited image Ie based on a textual instruction T: Ie = U(Ii; T) <ref:2503.20318#pg2>. This problem is framed similarly to the representation learning of images and text in CLIP, aiming to learn the semantics of edits within the CLIP space by encoding how reference images are transformed into their edited counterparts in relation to the provided instruction <ref:2503.20318#pg2>.
EditCLIP Pre-Training
The model modifies a standard CLIP visual encoder Fθ to accept a composite input image, where both the input and edited images are concatenated along the channel dimension, producing an edit embedding E = F˜θ(concat(Ii, Ie)) <ref:2503.20318#pg3>. The text encoder Gϕ encodes the editing instruction T into a textual embedding T = Gϕ(T) <ref:2503.20318#pg3>. The visual encoder is trained using a contrastive learning paradigm where only the visual encoder is trained while keeping the pre-trained text encoder frozen, aligning the learned editing space with the textual space using triplets consisting of an input image Ii, its edited counterpart Ie, and the corresponding editing instruction T <ref:2503.20318#pg3>.
EditCLIP for Exemplar-Based Image Editing
EditCLIP embeddings are used as a substitute for textual editing instructions in diffusion models like InstructPix2Pix (IP2P) <ref:2503.20318#pg4>. To train IP2P with EditCLIP, the edit embedding E is obtained from the input image Ii and its edited counterpart Ie using Equation (2), and this embedding is processed through a trainable linear layer followed by Layer Normalization to align it with the textual space originally used to train the diffusion model <ref:2503.20318#pg4>. During inference for a new query image Iq, the model is conditioned on the EditCLIP embedding produced from the exemplar image pair Ii and Ie, effectively modifying Equation (1) to perform exemplar-based image editing by generating an output image Io = U(Iq; F˜θ(concat(Ii, Ie))) <ref:2503.20318#pg4>.
EditCLIP for Evaluating Edits
The method provides two key metrics for automated evaluation. The EditCLIP-to-Text (EC2T) similarity metric quantifies how the input image transforms into the edited image and whether the changes align with the specified editing instructions, defined as EC2T(Ii, Ie, T) = cos(F˜θ(concat(Ii, Ie)), T) <ref:2503.20318#pg6>. For exemplar-based metrics, EditCLIP-to-EditCLIP (EC2EC) is computed as EC2EC(Ii, Ie, Iq, Io) = cos(F˜θ(concat(Ii, Ie)) F˜θ(concat(Iq, Io))) <ref:2503.20318#pg6>. These proposed metrics are shown to align more closely with human judgments than existing CLIP-based metrics <ref:2503.20318#pg2> and achieve the highest correlation with human evaluation for both text-based and exemplar-based editing tasks <ref:2503.20318#pg7>.
Quantitative Results
Experiments demonstrate the effectiveness of EditCLIP on two tasks, achieving state-of-the-art exemplar-based image editing with no computational overhead <ref:2503.20318#pg2>. The EC2EC metric performs the best among exemplarbased approaches, which is confirmed by a user study where its winning rate was larger than 50% against all baselines <ref:2503.20318#pg7>. Furthermore, EditCLIP embeddings are shown to be more reliable metrics for automated evaluation of both instruction-based and exemplar-based image editing methods due to their high Pearson correlation with human judgment <ref:2503.20318#pg7>.
Limitations and Future Work
Current training is solely on the IP2P dataset [4], which lacks edits like removal and deformation, suggesting that expanding training data could improve the quality and diversity of the embedding space <ref:2503.20318#pg2>. Future work includes applying EditCLIP to downstream tasks such as instruction caption generation, query-based editing pair retrieval, and extensions to video and 3D editing <ref:2503.20318#pg2>. Further improvements might involve exploring advanced training strategies like refined loss functions [51] or incorporating masks as an extra channel to enhance control over edit regions <ref:2503.20318#pg2>.
Conclusion
EditCLIP is a representation-learning approach that captures how images transform during edits, achieving state-of-the-art exemplar-based image editing with no computational overhead and serving as a reliable metric for evaluating edit quality and faithfulness to the reference image by aligning closely with human judgment. Such a metric can accelerate the development of image editing approaches by providing an evaluation metric that aligns better with human judgment compared to existing metrics <ref:2503.20318#pg2>.
REFERENCES
[1] Yuval Alaluf, Daniel Garibi, Or Patashnik, Hadar AverbuchElor, and Daniel Cohen-Or. Cross-image attention for zeroshot appearance transfer. In ACM SIGGRAPH 2024 Conference Papers, pages 1–12, 2024. <ref:2503.20318#pg10>
[2] Manuel Brack, Felix Friedrich, Dominik Hintersdorf, Lukas Struppek, Patrick Schramowski, and Kristian Kersting. Sega: Instructing diffusion using semantic dimensions. NeurIPS, 2023. <ref:2503.20318#pg11>
[3] Manuel Brack, Felix Friedrich, Katharina Kornmeier, Linoy Tsaban, Patrick Schramowski, Kristian Kersting, and Apolinaros Passos. Ledits++: Limitless image editing using textto-image models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024. <ref:2503.20318#pg10>
[4] Tim Brooks, Aleksander Holynski, and Alexei A Efros. Instructpix2pix: Learning to follow image editing instructions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 18392–18402, 2023. <ref:2503.20318#pg10>
[5] Mingdeng Cao, Xintao Wang, Zhongang Qi, Ying Shan, Xiaohu Qie, and Yinqiang Zheng. Masactrl: Tuning-free mutual self-attention control for consistent image synthesis and editing. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 22560–22570, 2023. <ref:2503.20318#pg11>
[6] Hila Chefer, Shir Gur, and Lior Wolf. Generic attentionmodel explainability for interpreting bi-modal and encoderdecoder transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 397–406, 2021. <ref:2503.20318#pg11>
[7] Ta-Ying Cheng, Prafull Sharma, Andrew Markham, Niki Trigoni, and Varun Jampani. Zest: Zero-shot material transfer from a single image. ECCV, 2024. <ref:2503.20318#pg10>
[8] Jooyoung Choi, Yunjey Choi, Yunji Kim, Junho Kim, and Sungroh Yoon. Custom-edit: Text-guided image editing with customized diffusion models. arXiv preprint arXiv:2305.15779, 2023. <ref:2503.20318#pg11>
[9] Jiwoo Chung, Sangeek Hyun, and Jae-Pil Heo. Style injection in diffusion: A training-free approach for adapting largescale diffusion models for style transfer. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 8795–8805, 2024. <ref:2503.20318#pg11>
[10] Aleksandar Cvejic, Abdelrahman Eldesokey, and Peter Wonka. Partedit: Fine-grained image editing using pretrained diffusion models. arXiv preprint arXiv:2502.04050, 2025. <ref:2503.20318#pg10>
[11] Wenkai Dong, Song Xue, Xiaoyue Duan, and Shumin Han. Prompt tuning inversion for text-driven image editing using diffusion models. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 7430–7440, 2023. <ref:2503.20318#pg11>
[12] Patrick Esser, Sumith Kulal, A. Blattmann, Rahim Entezari, Jonas Muller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, Dustin Podell, Tim Dockhorn, Zion English et al. Scaling rectified flow transformers for high-resolution image synthesis.
Improvements for AI systems
-
Improve exemplar-based image editing by replacing
text-based instructions in InstructPix2Pix [4] with EditCLIP embeddings computed from a reference exemplar image pair.
This allows forcomplex and precise edits, where describing the edit in natural language is challenging,
enabling users to apply multiple edits seamlessly based on a single reference pair. -
Enable automated evaluation of image editing pipelines by measuring
the similarity between the EditCLIP embedding of a given image pair and either a textual editing instruction or the EditCLIP embedding of another reference image pair.
This providesa reliable measure of edit quality and structural preservation,
aligning better with human judgments than existing CLIP-based metrics. -
Develop an exemplar-based editing system that can perform complex, multi-edit transfers in a single shot, as demonstrated by the result where
our method successfully transfers multiple edits from the exemplar in just one shot, while both IP2P and InstaManip fail.
This capability allows users to capture and transfer edits that aredifficult to express in natural language.
-
Enhance edit quality preservation by using learned hidden states instead of projected embeddings for conditioning, as suggested by the finding that
using hidden states from the last transformer layer before going to the CLIP projection layer is more effective to transfer the edit while preserving the input layout.
This directly addresses concerns where models mightblend the two images uncontrollably
when using other conditioning methods. -
Improve control over edit fidelity through dynamic guidance scales, as shown by Equation (14), which allows for two separate guidance scales:
edit guidance scale sE controls how the output image follows the edits, and image guidance scale sI controls how the output image resembles the input image.
This enables precise tuning of trade-offs between applying strong edits and preserving original structure.
Abstract
We introduce EditCLIP, a novel representation-learning approach for image editing. Our method learns a unified representation of edits by jointly encoding an input image and its edited counterpart, effectively capturing their transformation. To evaluate its effectiveness, we employ EditCLIP to solve two tasks: exemplar-based image editing and automated edit evaluation. In exemplar-based image editing, we replace text-based instructions in InstructPix2Pix with EditCLIP embeddings computed from a reference exemplar image pair. Experiments demonstrate that our approach outperforms state-of-the-art methods while being more efficient and versatile. For automated evaluation, EditCLIP assesses image edits by measuring the similarity between the EditCLIP embedding of a given image pair and either a textual editing instruction or the EditCLIP embedding of another reference image pair. Experiments show that EditCLIP aligns more closely with human judgments than existing CLIP-based metrics, providing a reliable measure of edit quality and structural preservation.
Sources
- Custom-Edit: Text-Guided Image Editing with Customized Diffusion Models
- PartEdit: Fine-Grained Image Editing using Pre-Trained Diffusion Models
- Planting a SEED of Vision in Large Language Model
- Improving Tuning-Free Real Image Editing with Proximal Guidance
- Prompt-to-Prompt Image Editing with Cross Attention Control
- FlowEdit: Inversion-Free Text-Based Editing Using Pre-Trained Flow Models
- Unleashing In-context Learning of Autoregressive Models for Few-shot Image Manipulation
- Hierarchical Text-Conditional Image Generation with CLIP Latents
- ReEdit: Multimodal Exemplar-Based Image Editing with Diffusion Models
- InstructBrush: Learning Attention-based Instruction Optimization for Image Editing
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models