Rethinking Text-Based Image Retrieval in Specific Domain

arXiv:2608.10524 · cs.CV, cs.AI · Submitted 2026-08-11 · Read on arXiv

Jingyang Tan, Sheng Yang, Yuanpeng Chen, Jian Wang, Nianjin Ye, Chen Xing, Lanpeng Jia

Harbin Institute of Technology · Fudan University · Changhong Intelligent Robot

cs.CV, cs.AI

Submitted: 2026-08-11

Updated: 2026-08-12

Comments: 13 pages

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 92/100

The gist: The paper introduces SecMM-TBIR, a multi-match benchmark for surveillance scenarios, and proposes the Semantic-Aware Fine-Tuning (SAFT) framework to address performance degradation in domain-specific

Terminology

Summary

The paper introduces SecMM-TBIR, a multi-match benchmark for surveillance scenarios, and proposes the Semantic-Aware Fine-Tuning (SAFT) framework to address performance degradation in domain-specific text-based image retrieval (TBIR).

The authors note that existing TBIR benchmarks are predominantly constructed on an exclusive single-match assumption between query and images, which fails to reflect practical system performance in specific domains (e.g., surveillance), where a single query often corresponds to multiple relevant candidate images. To address this, they design the Domain-Specific Multi-Match Text-based Image Retrieval (DSMM-TBIR) data engine, which leverages LLMs and VLMs with verification from multiple expert embedding models. Using this engine, they construct SecMM-TBIR, a benchmark comprising 50k surveillance images with 200 comprehensive queries across pedestrian and vehicle domains.

The paper also observes that vanilla contrastive learning in specific domains suffers from severe false negatives, forcing the model to push apart semantically similar pairs and thus degrading retrieval performance. To mitigate this, they propose SAFT, which incorporates two components: Semantic-Aware Soft-Label Supervision (SASS) and Intra-modal Structural Distillation (ISD). SASS accommodates potential positive text-image matches via a soft cross-modal alignment distribution, while ISD preserves relative correlations among images within the visual modality. The framework uses a frozen universal multi-modal embedding model (UniME-V2) as a teacher to provide cross-modal soft targets and image-to-image structural targets.

Experiments across diverse CLIP-like models (TinyCLIP, MobileCLIP, OpenCLIP) demonstrate that SAFT yields an average mAP@20 gain of 7.8 points on SecMM-TBIR over standard image-text contrastive (ITC) fine-tuning, while also improving general-domain performance. Specifically, compared with ITC, SAFT improves average mAP@20 by 5.4 and 10.3 points on pedestrian and vehicle subsets, respectively, and surpasses the CUSA baseline by 5.6 points on pedestrian and 10.7 points on vehicle retrieval. The framework also shows consistent improvements on general benchmarks (Flickr30K and MS-COCO), Fashion200K, and the ARO compositional reasoning benchmark.

Ablation studies reveal that fine-tuning the text encoder consistently degrades performance in specific domains, standard image self-supervision causes performance drops, and hard negative mining yields only marginal gains or can degrade retrieval performance due to rigid decision boundaries. The paper concludes that SAFT establishes a robust training and evaluation pipeline for domain-specific TBIR and enhances general-domain representation capabilities.

Improvements for AI systems

Improvements to AI Systems:

  1. Multi-Match Retrieval Capability: Replace single-match assumptions with a soft-label cross-modal alignment distribution (SASS), enabling the system to rank multiple relevant images for a single query instead of forcing a single correct answer. This directly improves recall in real-world domains like surveillance, where one query (e.g., person in red jacket near a blue car) has many valid matches.

  2. False-Negative Robustness in Contrastive Learning: Integrate semantic-aware soft labels into the contrastive loss, so the system no longer penalizes semantically similar but non-identical pairs as hard negatives. This prevents performance degradation from pushing apart legitimate matches, improving mAP@20 by 7.8 points on average over standard ITC fine-tuning.

  3. Intra-Modal Structural Preservation: Add Intra-modal Structural Distillation (ISD) to maintain relative correlations among images within the visual modality. This prevents the model from distorting image feature geometry during domain adaptation, preserving general-domain performance (e.g., +5.4 and +10.3 mAP@20 on pedestrian and vehicle subsets, respectively, while also improving Flickr30K and MS-COCO).

  4. Teacher-Guided Domain Adaptation: Use a frozen universal multimodal embedding model (UniME-V2) as a teacher to provide cross-modal soft targets and image-to-image structural targets. This allows the system to fine-tune on domain-specific data without catastrophic forgetting, achieving consistent gains on general benchmarks (Fashion200K, ARO) alongside domain-specific gains.

  5. Avoidance of Text-Encoder Fine-Tuning: Automatically freeze the text encoder during domain-specific fine-tuning, as ablation shows that fine-tuning it degrades performance. The improved system instead updates only the visual encoder and alignment layers, preserving semantic text understanding while adapting visual features.

  6. Benchmark-Driven Evaluation: Adopt the DSMM-TBIR data engine to generate multi-match benchmarks (e.g., SecMM-TBIR with 50k images, 200 queries) using LLM/VLM verification with multiple expert embeddings. This enables the system to be evaluated and iteratively improved on realistic, multi-match scenarios rather than single-match datasets.

What the Improved AI System Can Do:

  • Given a natural language query like person in a black hoodie carrying a backpack near a white van, it returns a ranked list of all relevant surveillance images (not just one), with higher precision and recall than standard CLIP-style models.

  • It maintains high accuracy on general image-text retrieval tasks (e.g., COCO captions) even after being fine-tuned for surveillance, avoiding the typical trade-off.

  • It is robust to ambiguous queries where multiple images are equally valid, because it learns a soft alignment distribution instead of a rigid one-hot target.

  • It can be deployed in real-time surveillance systems, where it improves vehicle and pedestrian search accuracy by 5.6–10.7 mAP@20 points over existing baselines, while also supporting compositional reasoning (e.g., red car left of a blue truck) without additional training.

Sources

Related papers