Heterogeneous Vision-Language Ensemble with Disagreement-Aware Reranking for Text-Based Person Anomaly Retrieval
Huu-An Vu, Cam Tu Tran Thi, Thanh Toan Le Ngo, Hoang Vo, Do Trung Hieu, Hieu Dinh Trung Pham, Khang Minh Le, Huy Minh Nhat Nguyen
Hanoi University of Science and Technology · University of Information Technology, VNU-HCM · Vietnam National University, Ho Chi Minh City · Ho Chi Minh City University of Science · VinUniversity · Vietnamese-German University · GenAI4E Lab
cs.CV, cs.AI
Submitted: 2026-08-13
Updated: 2026-08-14
Comments: Accepted at the ECCV 2026 Workshop
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 95/100
The gist: The paper "Heterogeneous Vision-Language Ensemble with Disagreement-Aware Reranking for Text-Based Person Anomaly Retrieval" presents the GENAI4E team's solution to AI City Challenge 2026 Track 4,
Terminology
Summary
The paper Heterogeneous Vision-Language Ensemble with Disagreement-Aware Reranking for Text-Based Person Anomaly Retrieval
presents the GENAI4E team's solution to AI City Challenge 2026 Track 4, addressing the task of text-based person anomaly retrieval (TBAPS). The task involves retrieving pedestrians exhibiting anomalous behaviors from a large image gallery using natural language descriptions, requiring fine-grained reasoning over pedestrian appearance, behaviors, object interactions, and scene context.
The proposed framework builds upon the Hybrid, Unified, and Iterative (HUI) baseline framework of Nguyen et al., which includes: "(1) the Local-Global Hybrid Perspective (LHP) module that enriches visual representations using probabilistic local and global image transformations; (2) Unified Image-Text (UIT) modeling, jointly optimizing image-text contrastive (ITC), image-text matching (ITM), masked language modeling (MLM), and masked image modeling (MIM) objectives; and (3) an iterative ensemble strategy that progressively refines retrieval predictions across multiple models."
The authors extend this baseline by integrating diverse pre-trained vision-language embedding models, including Voyage Multimodal, BGE-VL-v1.5-mmeb, and Qwen3-VL-Embedding-8B, into the iterative ensemble framework.
Since these models independently process the gallery and produce similarity matrices under different image orderings,
the authors introduce a score alignment procedure that projects all predictions onto a unified gallery indexing, enabling valid score-level fusion.
Additionally, the framework proposes "an Ensemble Disagreement routing mechanism that identifies ambiguous queries based on inter-model disagreement and selectively invokes a computationally expensive VLM cross-encoder reranker (gemini-3.1-flash-lite) only when necessary. The reranked predictions are
subsequently integrated with the original retrieval results via Reciprocal Rank Fusion (RRF)."
The paper's contributions are summarized as: (1) adopting the TBAPS baseline with LHP and UIT; (2) enriching the iterative ensemble with complementary pre-trained vision-language embedding models; (3) proposing a score alignment strategy for heterogeneous gallery orderings; and (4) introducing an Ensemble Disagreement routing mechanism with VLM reranking and RRF.
The evaluation is conducted on the official Pedestrian Anomaly Behavior (PAB) benchmark, which contains over one million synthetic image–text training pairs covering 1,000 routine and 1,600 anomalous action categories
and an evaluation set with 1,978 natural-language queries and 36,773 gallery images, including 34,795 distractors.
In ablation studies, the authors found that placing Qwen3-VL-Embedding as the final refinement stage
required careful weight tuning, with a minor adjustment to the weight schedule (from 0.9:0.92 to 0.88:0.9) results in a significant +1.29% mAP boost (from 88.96% to 90.25%).
They also found that incorporating BGE-VL-v1.5-mmeb yields the highest performance on the challenge setting,
and that stronger standalone embedding models do not necessarily yield the best ensemble performance
since the overall improvement largely depends on the complementarity between embedding spaces.
For the disagreement-aware reranking, the authors adopted a zero-tolerance disagreement threshold: a query is routed to the VLM if there is any inconsistency among the top-1 predictions of the ensemble configurations,
which isolates 299 ambiguous queries (15.1% of the test set).
This approach yields an 84.9% reduction in inference latency compared to exhaustive reranking.
The final system achieves 90.92% mAP, 85.13% Recall@1, 97.72% Recall@5, and 98.68% Recall@10
on the official AI City Challenge 2026 Track 4 benchmark. The authors conclude that combining complementary embedding spaces with selective vision-language reasoning provides an effective solution for large-scale text-based person anomaly retrieval under challenging surveillance scenarios.
Improvements for AI systems
Improvements to AI Systems:
-
Adaptive Inference Routing via Disagreement Detection: Implement a meta-level
disagreement gate
that monitors top-1 prediction consistency across an ensemble of diverse embedding models. When models disagree, the system dynamically invokes a high-cost, high-accuracy cross-encoder reranker; when they agree, it skips reranking entirely. This reduces inference latency by 85% while preserving accuracy, enabling real-time deployment on edge devices or in high-throughput surveillance pipelines. -
Heterogeneous Score Alignment for Ensemble Fusion: Develop a normalization and re-indexing layer that projects similarity scores from multiple pre-trained VLMs (with different internal gallery orderings and score distributions) onto a unified reference space. This allows seamless score-level fusion without retraining, improving robustness to model-specific biases and enabling plug-and-play integration of future embedding models.
-
Complementarity-Aware Model Selection: Instead of choosing the strongest single embedding model, the system automatically evaluates pairwise complementarity (e.g., via rank correlation or disagreement rate on a validation set) to select ensemble members. This prevents performance plateaus caused by redundant models and maximizes diversity, yielding higher mAP than any individual model.
-
Fine-Grained Weight Scheduling for Iterative Refinement: Introduce a dynamic weight-tuning mechanism for late-stage refinement models (e.g., large VLMs) that adjusts fusion weights based on query difficulty or iteration step. Small, systematic weight perturbations (e.g., 0.9→0.88) can yield >1% mAP gains, suggesting an automated hyperparameter optimizer (e.g., Bayesian search) for weight schedules.
-
Zero-Tolerance Ambiguity Flagging for Human-in-the-Loop: The disagreement threshold can be repurposed to flag ambiguous queries (e.g., 15% of cases) for human review or additional context retrieval, improving trust and accuracy in high-stakes anomaly detection scenarios (e.g., security alerts).
What the Improved AI System Can Do:
-
Process millions of gallery images with natural-language queries in near-real-time, using selective reranking only for genuinely ambiguous cases.
-
Achieve >90% mAP on text-based person anomaly retrieval, outperforming any single VLM, while cutting computational cost by an order of magnitude.
-
Seamlessly incorporate new pre-trained models without retraining, thanks to unified score alignment.
-
Automatically identify and escalate uncertain queries to human operators or more powerful reasoning models, reducing false negatives in surveillance.
-
Generalize to other fine-grained retrieval tasks (e.g., vehicle anomaly, product search) where heterogeneous embeddings and selective reasoning are beneficial.
Sources
- Qwen3-VL Technical Report
- RaSa: Relation and Sensitivity Aware Representation Learning for Text-based Person Search
- AnomalyLMM: Bridging Generative Knowledge and Discriminative Retrieval for Text-Based Person Anomaly Search
- Data Augmentation for Text-based Person Retrieval Using Large Language Models
- ESC: Emotional Self-Correction for Reliable Vision-Language Models
- ITSELF: Attention Guided Fine-Grained Alignment for Vision-Language Retrieval
- Passage Re-ranking with BERT
- Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks
- Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual Concepts
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models