The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric".
Jane: The paper was written by Sheng-Yu Wang, Yotam Nitzan, Aaron Hertzmann, Jun-Yan Zhu, Eli Shechtman et al. from Carnegie Mellon University and Adobe Research and University of California, Berkeley.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Title and Authors: Tom: Welcome back to the show, everyone. Today we’re digging into a paper that’s been making the rounds—it’s called “The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric.” And Jane, I gotta say, the title alone got me hooked.
Jane: Same here, Tom. It’s from a big team—Carnegie Mellon, Adobe Research, UC Berkeley. Sheng-Yu Wang, Yotam Nitzan, Aaron Hertzmann, Jun-Yan Zhu, Eli Shechtman, Alexei Efros, and Richard Zhang. That’s a who’s who of perceptual metrics and vision-language research.
Tom: Right, and the core idea is something we all kind of know but never really solved. When you look at two images, “similar” can mean a hundred different things. Same color, same pose, same background, same lighting—those are all different senses of similarity.
Jane: Exactly. And the paper points out that existing metrics like LPIPS or DreamSim just give you one number. They collapse all those senses into a single scalar. So if you ask “which image is more similar to this cow statue?”—you get an answer, but you have no idea why.
Tom: And that’s where the text-prompted part comes in. You can literally type “subject color” or “pose” or “ground surface,” and the metric gives you a similarity score conditioned on that aspect. It’s like asking a human to focus on one specific thing.
Jane: Yeah, and the authors built a whole dataset to train this. Over a million human similarity judgments on twenty-five thousand image triplets, each annotated across multiple free-form aspects. That’s a lot of careful annotation work.
Tom: And the implications are huge. Think about image retrieval, evaluating generative models, even understanding how humans perceive images. This could change how we build and evaluate vision systems.
Jane: Totally. And I love that they’re not just claiming it works—they benchmarked it against a ton of frontier models, including GPT and Gemini, and showed there’s a real gap. That’s the kind of honesty we need in this field.
Tom: So the title really captures it—there are many senses of visual similarity, and this paper is the first to take that seriously in a scalable, trainable way. I’m excited to see how they actually built the dataset and the model.
Jane: Me too. Let’s get into the summary and see what they found.
Summary: Tom: So Jane, we’re back with “The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric.” Let’s talk about what they actually did, because the dataset construction is pretty clever.
Jane: It really is. They used text-to-image models to generate triplets of images that differ along multiple aspects at once. So you get three images of, say, a cow statue, but one has different lighting, another has a different pose, another has a different background. And then they ask humans to pick the odd one out for each aspect.
Tom: And they didn’t just rely on the prompts. They used a vision-language model to refine the aspect list, dropping ones that aren’t visible and adding ones that are. That’s smart because text-to-image models don’t always follow prompts faithfully.
Jane: Right, and they also collected a second dataset from real vision algorithms—image editing, compositing, novel view synthesis, single-image three dee. That’s their out-of-distribution test set. So they’re not just testing on synthetic data they generated.
Tom: And the results? They benchmarked a ton of models—CLIP-based ones, VLM2Vec, Qwen3-VL-Embedding, GPT-five point four, Gemini three point one Pro. And the gap to human consensus was pretty big. Human consensus on their odd-one-out task was around sixty-seven point five percent, and the best baseline was Gemini at fifty-eight point four percent.
Jane: That’s a real gap. And their model, TPIPS, gets to sixty-four point seven percent on the odd-one-out task, which is much closer to human agreement. On the out-of-distribution 2AFC task, they get seventy-seven point nine percent versus the best baseline at seventy-three point nine percent. So it generalizes beyond the training distribution.
Tom: And they tested three architectures—late fusion, mid fusion, early fusion. Early fusion, where the two images interact from the start, does the best. But late fusion is almost as good and way more efficient because you can precompute embeddings.
Jane: That efficiency matters for real-world use. And they also showed the metric works for retrieval—same query image, different aspects, different nearest neighbors. That’s a capability we just didn’t have before.
Tom: It’s like having a search engine where you can say “find me images that are similar in color but not in composition.” That’s powerful.
Jane: And they even did compositional retrieval—combining multiple queries with different aspects. Like “a painting of two people” plus “Monet’s brushwork” gives you Impressionist paintings of two people. That’s wild.
Tom: It really is. But I’m curious about the improvements they suggest. What’s next for this line of work?
Improvements: Tom: So we’re still on “The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric.” Jane, what improvements do the authors suggest for future work?
Jane: Well, they’re pretty honest about limitations. The metric is expensive—it’s a VLM-based model, so it’s not as cheap as LPIPS. And the embedding-based retrieval requires separate indexing for each aspect condition, which is a practical bottleneck.
Tom: Yeah, that’s a real engineering concern. But they also mention data diversity. Their training data comes from FLUX-Reason-6M prompts, which are more aesthetically curated. So the metric might not generalize perfectly to all types of images.
Jane: And the aspect proposals are VLM-generated, so they might miss aspects that VLMs can’t capture. That’s a systematic blind spot. But they suggest more diverse image sources and human-annotated aspects for future work.
Tom: I think the biggest improvement they hint at is using the metric as feedback for generative models. You could train an image editing model or a text-to-image model to better match human perception on specific aspects.
Jane: Exactly. And they mention interpretability—using the metric to discover what features neural networks are encoding. You could ask “which visual properties does this neuron respond to?” and get a human-aligned answer.
Tom: That’s a big deal for safety and transparency. If we can automatically characterize what models are doing, we can better audit them.
Jane: And they also suggest that as VLMs improve, their pipeline naturally benefits. So this isn’t a one-off—it’s a foundation that gets better with better base models.
Tom: So the improvements are really about scaling up data, making it more efficient, and applying it to real-world tasks like model evaluation and interpretability.
Jane: Right. And I think the compositional retrieval part is under-explored. They showed it works, but there’s room to make it more robust and useful for actual search systems.
Tom: Agreed. So what’s the big picture here? Where does this leave us?
Conclusion: Tom: Alright, we’re wrapping up our discussion on “The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric.” Jane, what’s the takeaway for our listeners?
Jane: I think the biggest takeaway is that similarity isn’t one thing. It’s many things, and this paper gives us a way to ask which sense of similarity we care about. That’s a fundamental shift in how we think about perceptual metrics.
Tom: And they backed it up with a massive dataset, rigorous benchmarking, and a model that actually beats the frontier VLMs on their own terms. That’s no small feat.
Jane: The implications go beyond just measuring similarity. This could change how we evaluate generative models, how we build retrieval systems, and even how we understand what neural networks are doing internally.
Tom: And the authors are sharing their code, data, and trained models. So this isn’t just a paper—it’s a resource the whole community can build on.
Jane: Exactly. And I think the most exciting part is that this is just the beginning. As VLMs get better, this pipeline gets better. And as we collect more diverse data, the metric becomes more robust.
Tom: So for anyone working in vision, this is a paper you need to read. It’s not just an incremental step—it’s a new way of thinking about a problem we thought we’d solved.
Jane: And with that, we’re saying goodbye to “The Many Senses of Visual Similarity.” Next up, we’ve got a paper on something completely different, so stay tuned.
Tom: Thanks for listening, everyone. See you next time.
Sheng-Yu Wang, Yotam Nitzan, Aaron Hertzmann, Jun-Yan Zhu, Eli Shechtman, Alexei A. Efros, Richard Zhang
Carnegie Mellon University · Adobe Research · University of California, Berkeley
cs.CV, cs.LG
Submitted: 2026-08-17
Updated: 2026-08-18
Comments: Project Webpage: https://peterwang512.github.io/TPIPS
Code: https://github.com/QwenLM/Qwen3.5
Project page: https://peterwang512.github.io/TPIPS
License: http://creativecommons.org/licenses/by-sa/4.0/
Importance score: 71/100
The gist: The paper introduces a new approach to visual similarity that is context-dependent, addressing the limitation of existing perceptual metrics that collapse multiple senses of similarity into a single
Key concepts
- Many Senses of Visual Similarity
- This concept refers to the idea that two images can be similar in many different ways, such as color, pose, background, or lighting. Existing metrics often collapse these many senses into one single score.
- Text-Prompted Image Perceptual Metric
- This is a new metric where users can type specific visual aspects—like "subject color" or "pose"—to get a similarity score conditioned on that aspect, allowing for targeted comparisons between images.
- Compositional Retrieval
- This capability allows users to combine multiple queries with different aspects to find highly specific results. For example, searching for a painting of two people and specifying 'Monet’s brushwork' yields Impressionist paintings of two people.
Terminology
Summary
The paper introduces a new approach to visual similarity that is context-dependent, addressing the limitation of existing perceptual metrics that collapse multiple senses of similarity into a single scalar value.
The authors state: "Human visual similarity judgments are context-dependent. For example, two images may be similar in shape but distinct in color. Existing perceptual similarity metrics, however, collapse these nuances into a single scalar value, offering no mechanism to condition on specific aspects."
To bridge this gap, they collect a large-scale dataset: we collect a large-scale dataset of one million human similarity judgments over 25K image triplets.
They extend the standard triplet protocol: "we extend this protocol to perceptual judgements conditioned on various visual aspects. We collect 260K triplet-aspect condition combinations, where each triplet is annotated conditioned on multiple free-form visual aspects, each aspect yielding a separate set of similarity judgments. The triplets are generated using text-to-image models and are designed to be challenging, targeting a middle ground where
images that differ along multiple aspects simultaneously — through fine-grained variations in properties such as color, lighting, and texture."
They also collect a second, out-of-distribution dataset: we collect a second set of human similarity judgments over outputs generated by external algorithms, designed for four different vision and graphics tasks, spanning both generative and non-generative pipelines.
This uses a 2-alternative-forced-choice (2AFC) protocol.
Benchmarking reveals a performance gap: Benchmarking a broad range of frontier vision-language models (VLMs) reveals a considerable performance gap compared to human annotators’ concensus.
They find that even the strongest current VLM systems often do not agree with human similarity judgments on specified aspects.
They propose a new model: Leveraging our data, we fine-tune a VLM to produce our Text-Prompted Image Perceptual Similarity (TPIPS) metric, capturing multiple senses of visual similarity depending on the specified text prompt.
They explore three architectures: late fusion (embeddings), mid fusion (activation distances), and early fusion. They note: In the controlled setting where we fine-tune all model types under the same base model, early fusion consistently outperforms both middle and late fusion, albeit by a modest margin.
Results show: Our method (red bars) consistently outperforms baselines in both in-domain odd-one-out tests and out-of-domain 2AFC tests.
Specifically, "Fine-tuning on our data substantially narrows the gap to the human consensus, reducing it from 9.1% to 2.8% on the odd-one-out task. Notably, these gains translate to the held-out 2AFC distribution, reducing error from 15.8% to 10.0%, above baselines. They also demonstrate generalization:
our model outperforms baselines in this setting too, suggesting that it can serve as a general-purpose perceptual metric for aspect-conditioned evaluation."
They demonstrate several downstream applications: "evaluation of generative vision models, where output and reference similarity is measured along specified visual axes, nearest-neighbor retrieval, where the same query image returns different neighbors under different aspects, and compositional retrieval, where multiple query images, each conditioned on a different visual aspect, can be combined to find images that match the criteria."
The paper concludes with a discussion of broader impacts and limitations. They note the metric can serve as a feedback signal for training models and aid in interpretability. Limitations include: Our VLM-based metric is substantially more expensive to compute than LPIPS or DreamSim,
our annotator pool is not representative of all viewers,
and our method is also limited by training data diversity.
They state: "as VLMs continue to improve, our pipeline naturally benefits — and we view our human perceptual data as a complementary signal toward building more human-aligned models, offering a new angle for collecting fine-grained human perceptual judgments."
Improvements for AI systems
Based on the paper, here are the specific improvements I can make to AI systems:
Improvement: Replace single-scalar similarity metrics (LPIPS, DreamSim, SSIM) with a text-prompted perceptual metric that conditions on specific visual aspects.
What the improved system can do:
-
Given a reference image and a text prompt like
lighting warmth,
object pose,
orbackground texture,
return a similarity score that specifically measures that aspect, not a global average. -
Resolve ambiguous similarity judgments: when two images are similar in shape but different in color, the system can be told which aspect to prioritize.
-
Provide a fine-grained evaluation signal for generative models: instead of asking
is this generated image good overall?
, askis the generated image's subject pose similar to the reference?
Abstract
Human visual similarity judgments are context-dependent. For example, two images may be similar in shape but distinct in color. Existing perceptual similarity metrics, however, collapse these nuances into a single scalar value, offering no mechanism to condition on specific aspects. To bridge this gap, we introduce a large-scale dataset of human similarity judgments over image triplets, where each triplet is annotated across multiple, free-form semantic aspects of similarity. Benchmarking a broad range of frontier vision-language models (VLMs) reveals a considerable performance gap compared to human annotators' consensus. Leveraging our data, we fine-tune a VLM to produce our Text-Prompted Image Perceptual Similarity (TPIPS) metric, capturing multiple senses of visual similarity depending on the specified text prompt. We demonstrate that TPIPS aligns more closely with human perception and generalizes reliably beyond the training distribution. Finally, we show that TPIPS unlocks new capabilities in text-guided retrieval, compositional search, and the fine-grained evaluation of generative models. Our code, data, and trained models are at https://peterwang512.github.io/TPIPS
Sources
- FocalLens: Instruction Tuning Enables Zero-Shot Conditional Image Representations
- Omni-Attribute: Open-vocabulary Attribute Encoder for Visual Concept Personalization
- Gemini: A Family of Highly Capable Multimodal Models
- Qwen3-VL Technical Report
- E-LPIPS: Robust Perceptual Image Similarity via Random Transformation Ensembles
- ImageReward: Learning and Evaluating Human Preferences for Text-to-Image Generation
- Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis
- Cooperative Embeddings for Instance, Attribute and Category Retrieval
- Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
- Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
- LLaVA-OneVision: Easy Visual Task Transfer
- VLM2Vec-V2: Advancing Multimodal Embedding for Videos, Images, and Visual Documents
- FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark
- Steerable Visual Representations
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- Dual-Process Image Generation
- DanceGRPO: Unleashing GRPO on Visual Generation
- One-Step is Enough: Sparse Autoencoders for Text-to-Image Diffusion Models
- Qwen3 Technical Report
- GEditBench v2: A Human-Aligned Benchmark for General Image Editing
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models