Image Recognition with Vision and Language Embeddings of VLMs

arXiv:2509.09311 · cs.CV · Submitted 2025-09-11 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Image Recognition with Vision and Language Embeddings of VLMs".

Jane: The gist The authors conduct a comprehensive evaluation of both language-guided and vision-only image classification with dual-encoder VLMs, showing that language and vision offer complementary strengths,

Tom: First, who's behind it and why it matters.

Title and authors: Tom: So let's talk about the title and who wrote this. "Image Recognition with Vision and Language Embeddings of VLMs." It tells you immediately that they are using both vision embeddings and language embeddings in their models for recognition.

Jane: And the authors are Volkov, Kisel, Janouskova, and Matas. They’re researchers at the Czech Technical University in Prague.

Lu: These authors are clearly digging into how these multimodal models actually perform on tasks like image classification when they’re using both visual and textual information simultaneously.

Meng: It sounds like their main goal is to give researchers a better way to understand which part of the vision-language model—the language side or the vision side—is doing what it’s doing best for different kinds of images.

Lalam: It points toward figuring out how to use these models more effectively, especially when we want strong zero-shot classification, which is when a model recognizes something it hasn't been explicitly trained on before.

The paper's summary: Tom: They summarize the work by saying they are conducting a comprehensive evaluation of both language-guided and vision-only image classification using these dual-encoder VLMs on the standard ImageNet-1k set and a cleaner version of it <ref:2509.09311#pg1,a comprehensive evaluation of both language-guided and>.

Jane: Basically, they set up this comparison to see how well each approach stacks up against each other, paying attention to things like prompt design and the size of the reference set used in their methods.

Lu: They specifically analyze key factors like prompt design, class diversity, how many neighbors are in a k-NN classifier, and how big that reference set needs to be for good results.

Meng: It’s interesting they are looking at these specific factors because it suggests that the best way to use a VLM isn't just about having the biggest model; it's about setting up the right conditions for classification.

Lalam: Their main finding in this section is showing that language and vision really do offer complementary strengths, meaning some classes are better when you give them text prompts, and others are better handled by looking at visual similarity.

The paper's improvements: Tom: Now let’s talk about what they suggest as improvements. They introduce a simple, learning-free fusion method that dynamically switches between vision and language predictions based on how precise each classifier is for a specific class.

Jane: That dynamic selection process is quite clever because it lets the model pick the prediction that has higher precision for that particular image and class combination.

Lu: They say this simple, parametric-free learning method results in an accuracy improvement on both validation and cleaner sets compared to using just vision-only or just language-based methods.

Meng: An accuracy improvement of zero point four percent is modest, but they claim it's still an actual gain when you combine them this way, which is something practical for engineers to consider.

Lalam: The best overall accuracy achieved by this combined approach in their testing was eighty-six point nine zero on the validation set and the cleaner set, which shows that even a simple non-parametric combination can help.

Conclusion: Tom: So to wrap things up, the authors conclude that some classes, like those with high visual diversity, actually benefit more from textual representations than others.

Jane: And they propose this simple method of dynamically selecting between vision and language predictions based on per-class precision as a way to get the best result overall.

Lu: It really shows that you don't have to do something overly complicated or train a massive model; you can use these dual-encoder VLMs in a smarter, combination way.

Meng: For practical applications, this means we might be able to build systems where we can intelligently choose whether to trust the language prompt or the visual feature map for a given input.

Lalam: So, overall, this paper on "Image Recognition with Vision and Language Embeddings of VLMs" demonstrates that language and vision offer complementary strengths, and their simple fusion method gives us a way to get better accuracy than using either modality alone.

Czech Technical University in Prague

cs.CV

Submitted: 2025-09-11

Updated: 2026-10-08

Code: https://github.com/gonikisgo/bmvc2025-vlm-image-recognition

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 80/100

The gist: The gist The authors conduct a comprehensive evaluation of both language-guided and vision-only image classification with dual-encoder VLMs, showing that language and vision offer complementary

Key concepts

Dual-encoder VLMs
These are AI models with separate encoders for images and language. They process an image and text separately before combining their information to make a classification decision. This architecture allows the model to leverage both visual data and textual descriptions simultaneously.
Language-Based Classification
This method classifies an image based on a text prompt provided by the user, using the VLM's language understanding capabilities. The performance depends heavily on how well the prompt is designed, such as choosing specific class names or using different template structures for querying the model.
Vision-Only Classification with k-NN
This approach classifies an image solely based on its visual features by comparing it to a stored reference set of labeled images. The k-Nearest Neighbors (k-NN) algorithm finds the most similar examples in this set to determine the class label, showing that vision alone can perform surprisingly well.
Complementary Fusion Method
This is a simple, learning-free technique that dynamically chooses between the predictions from the language model and the vision model for each class. It selects whichever prediction has higher precision for that specific class, resulting in a modest accuracy gain by combining the strengths of both modalities.

Terminology

Summary

The gist The authors conduct a comprehensive evaluation of both language-guided and vision-only image classification with dual-encoder VLMs, showing that language and vision offer complementary strengths, with some classes favouring textual prompts and others better handled by visual similarity

Evaluation Setup

The study focuses on dual-encoder VLMs [1, 7, 16, 22, 29, 30], which are architectures with separate image and language encoders The goal is to help researchers better navigate the rapidly evolving VLM landscape The models are benchmarked on the standard ImageNet-1k [17] dataset and a curated “Cleaner” variant [9] with higher ground-truth label accuracy created by reviewing and building on prior relabelling effort [2, 12, 18, 23>

Language-Based Classification Analysis

The performance comparison is conducted in a standard setup on the ImageNet-1k validation set The key factors affecting accuracy are analyzed, including prompt design, class diversity, the number of neighbours in k-NN, and reference set size The standard zero-shot VLM ImageNet-1k evaluation is based on the average prompt embedding (2) which is computed over seven hand-crafted templates Results show that the original CLIP [16] is significantly outperformed by all newer models, whose accuracy is in a fairly narrow range of 91-93% on the Cleaner and 83-85% on the original validation The best zero-shot method on the validation set is SigLIP 2 [22]

Text Prompting Techniques

The paper investigates the topic with results in Table 2, examining the influence of class names with WordNet, OpenAI [16], and OpenAI+ [9] class names The standard average prompt approach is denoted Avg, while Avg′ denotes the ensemble of seven standard prompts plus the no-context template Results indicate that OpenAI+ class names lead to higher recognition accuracy compared to WordNet class names Templates like a origami a class name and art of the a class name perform worse than the standard itap of a class name (see Table 2)

Vision-Only Classification with k-NN

The vision-only capability is investigated using the standard k-Nearest Neighbours (k-NN) algorithm, which is learning-free and only needs the storage of a labelled reference set The classifier assigns a label by identifying the most frequent class among the k nearest reference samples Results show that vision-only classification outperforms zero-shot language classification, matches the supervised results of EfficientNetV2, but remains worse than EfficientNet-L2 and EVA02 The optimal k falls in the same interval from 7 to 11 across models and dataset splits

Complementary Fusion Method

To exploit this complementarity, a simple, learning-free fusion method is introduced that dynamically selects between vision and language predictions based on classifiers per-class precision The approach involves a non-parametric learning phase followed by an inference step where the final prediction p is the one with higher precision corresponding to the predicted classes This method resulted in an accuracy improvement on both validation and Cleaner sets compared to vision-only and language-based results, which is a modest 0.4% The best overall accuracy achieved by this combined approach is 86.90

Conclusion

The analysis reveals that some classes, e.g., classes with high visual diversity, benefit more from textual representations, while others are better captured visually The authors propose a simple, learning-free combination method that outperforms either approach alone This demonstrates the power of even a simple non-parametric learning classifier combination for image recognition The final results show modest gains from simple non-parametric fusion

--- Page 1 ---

The gist The authors conduct a comprehensive evaluation of both language-guided and vision-only image classification with dual-encoder VLMs, showing that language and vision offer complementary strengths, with some classes favouring textual prompts and others better handled by visual similarity

Text Prompting Techniques

The paper investigates the topic with results in Table 2, examining the influence of class names with WordNet, OpenAI [16], and OpenAI+ [9] class names The standard average prompt approach is denoted Avg, while Avg′ denotes the ensemble of seven standard prompts plus the no-context template Results indicate that OpenAI+ class names lead to higher recognition accuracy compared to WordNet class names

Complementary Fusion Method

To exploit this complementarity, a simple, learning-free fusion method is introduced that dynamically selects between vision and language predictions based on classifiers per-class precision The approach involves a non-parametric learning phase followed by an inference step where the final prediction p is the one with higher precision corresponding to the predicted classes This method resulted in an accuracy improvement on both validation and Cleaner sets compared to vision-only

Improvements for AI systems

  1. textbf Language-Vision Fusion for Robust Classification (Table 5): The proposed Precision-Based combination strategy dynamically selects between vision-only and language-based predictions based on classifiers per-class precision. This system can perform image recognition by outputting the prediction corresponding to the classifier with higher precision, which is shown to yield an accuracy improvement on both validation and Cleaner sets compared to either approach alone.

  2. textbf Adaptive Prompt Engineering for Zero-Shot Accuracy (Table 6): The system can improve zero-shot classification by employing optimized class names, such as OpenAI+ class names, in conjunction with prompt ensembling (Avg' OpenAI+). This allows the AI to achieve higher validation accuracy by leveraging text representations that better align with image content and reduce language ambiguity.

  3. textbf Data-Driven Class Prioritization for k-NN Optimization (Figure 5): The system can optimize its vision-only classification performance by dynamically selecting the best 'k' value for k-NN based on class characteristics. This allows the AI to achieve accuracy improvements of up to 20% by recognizing that the best kNN settings vary between classes, specifically targeting classes with high ground truth label noise or fine-grained categories.

  4. textbf Error Mitigation via Oracle Scenarios (Figure 2): The system can utilize class-level and image-level oracle evaluations to select the most accurate classification strategy. For instance, the image-level oracle can push accuracy to 91.96%, demonstrating how selecting the best template for each individual image maximizes performance despite inherent dataset flaws like label noise or distribution shift.

Abstract

Vision-language models (VLMs) have enabled strong zero-shot classification through image-text alignment. Yet, their purely visual inference capabilities remain under-explored. In this work, we conduct a comprehensive evaluation of both language-guided and vision-only image classification with a diverse set of dual-encoder VLMs, including both well-established and recent models such as SigLIP 2 and RADIOv2.5. The performance is compared in a standard setup on the ImageNet-1k validation set and its label-corrected variant. The key factors affecting accuracy are analysed, including prompt design, class diversity, the number of neighbours in k-NN, and reference set size. We show that language and vision offer complementary strengths, with some classes favouring textual prompts and others better handled by visual similarity. To exploit this complementarity, we introduce a simple, learning-free fusion method based on per-class precision that improves classification performance. The code is available at: https://github.com/gonikisgo/bmvc2025-vlm-image-recognition.

Sources

Related papers