The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric

summary

Video file (mp4)

The gist

The paper introduces a new approach to visual similarity that is context-dependent, addressing the limitation of existing perceptual metrics that collapse multiple senses of similarity into a single

In short

The episode discusses a paper introducing "The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric." The paper addresses how similarity has many aspects, moving beyond single metrics like LPIPS. The authors built a dataset and metric that allows users to query similarity based on specific visual senses like color or pose, showing significant gaps in current models and offering new ways to evaluate generative models.

Key concepts

Many Senses of Visual Similarity
This concept refers to the idea that two images can be similar in many different ways, such as color, pose, background, or lighting. Existing metrics often collapse these many senses into one single score.
Text-Prompted Image Perceptual Metric
This is a new metric where users can type specific visual aspects—like "subject color" or "pose"—to get a similarity score conditioned on that aspect, allowing for targeted comparisons between images.
Compositional Retrieval
This capability allows users to combine multiple queries with different aspects to find highly specific results. For example, searching for a painting of two people and specifying 'Monet’s brushwork' yields Impressionist paintings of two people.

Terminology used across episodes

This episode discusses

The paper

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric · Read on arXiv

Sheng-Yu Wang, Yotam Nitzan, Aaron Hertzmann, Jun-Yan Zhu, Eli Shechtman, Alexei A. Efros, Richard Zhang

Carnegie Mellon University · Adobe Research · University of California, Berkeley

Human visual similarity judgments are context-dependent. For example, two images may be similar in shape but distinct in color. Existing perceptual similarity metrics, however, collapse these nuances into a single scalar value, offering no mechanism to condition on specific aspects. To bridge this gap, we introduce a large-scale dataset of human similarity judgments over image triplets, where each triplet is annotated across multiple, free-form semantic aspects of similarity. Benchmarking a broad range of frontier vision-language models (VLMs) reveals a considerable performance gap compared to human annotators' consensus. Leveraging our data, we fine-tune a VLM to produce our Text-Prompted Image Perceptual Similarity (TPIPS) metric, capturing multiple senses of visual similarity depending on the specified text prompt. We demonstrate that TPIPS aligns more closely with human perception and generalizes reliably beyond the training distribution. Finally, we show that TPIPS unlocks new capabilities in text-guided retrieval, compositional search, and the fine-grained evaluation of generative models. Our code, data, and trained models are at https://peterwang512.github.io/TPIPS

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric".

Jane: The paper was written by Sheng-Yu Wang, Yotam Nitzan, Aaron Hertzmann, Jun-Yan Zhu, Eli Shechtman et al. from Carnegie Mellon University and Adobe Research and University of California, Berkeley.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title and Authors: Tom: Welcome back to the show, everyone. Today we’re digging into a paper that’s been making the rounds—it’s called “The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric.” And Jane, I gotta say, the title alone got me hooked.

Jane: Same here, Tom. It’s from a big team—Carnegie Mellon, Adobe Research, UC Berkeley. Sheng-Yu Wang, Yotam Nitzan, Aaron Hertzmann, Jun-Yan Zhu, Eli Shechtman, Alexei Efros, and Richard Zhang. That’s a who’s who of perceptual metrics and vision-language research.

Tom: Right, and the core idea is something we all kind of know but never really solved. When you look at two images, “similar” can mean a hundred different things. Same color, same pose, same background, same lighting—those are all different senses of similarity.

Jane: Exactly. And the paper points out that existing metrics like LPIPS or DreamSim just give you one number. They collapse all those senses into a single scalar. So if you ask “which image is more similar to this cow statue?”—you get an answer, but you have no idea why.

Tom: And that’s where the text-prompted part comes in. You can literally type “subject color” or “pose” or “ground surface,” and the metric gives you a similarity score conditioned on that aspect. It’s like asking a human to focus on one specific thing.

Jane: Yeah, and the authors built a whole dataset to train this. Over a million human similarity judgments on twenty-five thousand image triplets, each annotated across multiple free-form aspects. That’s a lot of careful annotation work.

Tom: And the implications are huge. Think about image retrieval, evaluating generative models, even understanding how humans perceive images. This could change how we build and evaluate vision systems.

Jane: Totally. And I love that they’re not just claiming it works—they benchmarked it against a ton of frontier models, including GPT and Gemini, and showed there’s a real gap. That’s the kind of honesty we need in this field.

Tom: So the title really captures it—there are many senses of visual similarity, and this paper is the first to take that seriously in a scalable, trainable way. I’m excited to see how they actually built the dataset and the model.

Jane: Me too. Let’s get into the summary and see what they found.

Summary: Tom: So Jane, we’re back with “The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric.” Let’s talk about what they actually did, because the dataset construction is pretty clever.

Jane: It really is. They used text-to-image models to generate triplets of images that differ along multiple aspects at once. So you get three images of, say, a cow statue, but one has different lighting, another has a different pose, another has a different background. And then they ask humans to pick the odd one out for each aspect.

Tom: And they didn’t just rely on the prompts. They used a vision-language model to refine the aspect list, dropping ones that aren’t visible and adding ones that are. That’s smart because text-to-image models don’t always follow prompts faithfully.

Jane: Right, and they also collected a second dataset from real vision algorithms—image editing, compositing, novel view synthesis, single-image three dee. That’s their out-of-distribution test set. So they’re not just testing on synthetic data they generated.

Tom: And the results? They benchmarked a ton of models—CLIP-based ones, VLM2Vec, Qwen3-VL-Embedding, GPT-five point four, Gemini three point one Pro. And the gap to human consensus was pretty big. Human consensus on their odd-one-out task was around sixty-seven point five percent, and the best baseline was Gemini at fifty-eight point four percent.

Jane: That’s a real gap. And their model, TPIPS, gets to sixty-four point seven percent on the odd-one-out task, which is much closer to human agreement. On the out-of-distribution 2AFC task, they get seventy-seven point nine percent versus the best baseline at seventy-three point nine percent. So it generalizes beyond the training distribution.

Tom: And they tested three architectures—late fusion, mid fusion, early fusion. Early fusion, where the two images interact from the start, does the best. But late fusion is almost as good and way more efficient because you can precompute embeddings.

Jane: That efficiency matters for real-world use. And they also showed the metric works for retrieval—same query image, different aspects, different nearest neighbors. That’s a capability we just didn’t have before.

Tom: It’s like having a search engine where you can say “find me images that are similar in color but not in composition.” That’s powerful.

Jane: And they even did compositional retrieval—combining multiple queries with different aspects. Like “a painting of two people” plus “Monet’s brushwork” gives you Impressionist paintings of two people. That’s wild.

Tom: It really is. But I’m curious about the improvements they suggest. What’s next for this line of work?

Improvements: Tom: So we’re still on “The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric.” Jane, what improvements do the authors suggest for future work?

Jane: Well, they’re pretty honest about limitations. The metric is expensive—it’s a VLM-based model, so it’s not as cheap as LPIPS. And the embedding-based retrieval requires separate indexing for each aspect condition, which is a practical bottleneck.

Tom: Yeah, that’s a real engineering concern. But they also mention data diversity. Their training data comes from FLUX-Reason-6M prompts, which are more aesthetically curated. So the metric might not generalize perfectly to all types of images.

Jane: And the aspect proposals are VLM-generated, so they might miss aspects that VLMs can’t capture. That’s a systematic blind spot. But they suggest more diverse image sources and human-annotated aspects for future work.

Tom: I think the biggest improvement they hint at is using the metric as feedback for generative models. You could train an image editing model or a text-to-image model to better match human perception on specific aspects.

Jane: Exactly. And they mention interpretability—using the metric to discover what features neural networks are encoding. You could ask “which visual properties does this neuron respond to?” and get a human-aligned answer.

Tom: That’s a big deal for safety and transparency. If we can automatically characterize what models are doing, we can better audit them.

Jane: And they also suggest that as VLMs improve, their pipeline naturally benefits. So this isn’t a one-off—it’s a foundation that gets better with better base models.

Tom: So the improvements are really about scaling up data, making it more efficient, and applying it to real-world tasks like model evaluation and interpretability.

Jane: Right. And I think the compositional retrieval part is under-explored. They showed it works, but there’s room to make it more robust and useful for actual search systems.

Tom: Agreed. So what’s the big picture here? Where does this leave us?

Conclusion: Tom: Alright, we’re wrapping up our discussion on “The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric.” Jane, what’s the takeaway for our listeners?

Jane: I think the biggest takeaway is that similarity isn’t one thing. It’s many things, and this paper gives us a way to ask which sense of similarity we care about. That’s a fundamental shift in how we think about perceptual metrics.

Tom: And they backed it up with a massive dataset, rigorous benchmarking, and a model that actually beats the frontier VLMs on their own terms. That’s no small feat.

Jane: The implications go beyond just measuring similarity. This could change how we evaluate generative models, how we build retrieval systems, and even how we understand what neural networks are doing internally.

Tom: And the authors are sharing their code, data, and trained models. So this isn’t just a paper—it’s a resource the whole community can build on.

Jane: Exactly. And I think the most exciting part is that this is just the beginning. As VLMs get better, this pipeline gets better. And as we collect more diverse data, the metric becomes more robust.

Tom: So for anyone working in vision, this is a paper you need to read. It’s not just an incremental step—it’s a new way of thinking about a problem we thought we’d solved.

Jane: And with that, we’re saying goodbye to “The Many Senses of Visual Similarity.” Next up, we’ve got a paper on something completely different, so stay tuned.

Tom: Thanks for listening, everyone. See you next time.

More episodes

← Home