MLLM-Routed Heterogeneous Ensembles for Robust Cross-Dataset Image Classification

arXiv:2608.13463 · cs.CV, cs.AI, cs.CL, cs.LG · Submitted 2026-08-13 · Read on arXiv

University of Tennessee, Knoxville · The Bredesen Center for Interdisciplinary Research and Graduate Education

cs.CV, cs.AI, cs.CL, cs.LG

Submitted: 2026-08-13

Updated: 2026-08-27

Comments: 8 pages, 4 figures, 7 tables

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

The gist: The paper proposes ARMDIL, an Adaptive Router for Multi-Domain Image classification with LLMs.

Terminology

Summary

The paper proposes ARMDIL, an Adaptive Router for Multi-Domain Image classification with LLMs. ARMDIL is an ensemble that uses a multimodal large language model (MLLM) agent to dynamically route each image to the most suitable vision backbone. The diverse ensemble employs convolutional neural networks (ResNets), self-supervised representation learners (SSL), and vision-language models (VLMs), each trained on a unified label space constructed from multiple image datasets with differing distributions and characteristics.

The paper states: "We propose ARMDIL, an Adaptive Router for Multi-Domain Image classification with LLMs. ARMDIL is an ensemble that uses a multimodal large language model (MLLM) agent to dynamically route each image to the most suitable vision backbone. Our diverse ensemble employs convolutional neural networks (ResNets), self-supervised representation learners (SSL), and vision-language models (VLMs), each trained on a unified label space constructed from multiple image datasets with differing distributions and characteristics."

The framework is designed to address the challenge that Modern image classification models excel when trained on single task-specific datasets but often struggle to generalize across domains and difficulty levels. The paper notes that a model highly optimized for one task frequently fails on another.

The methodology involves four cross-domain datasets: CIFAR10 (general objects), FER2013 (facial expressions), EuroSAT (satellite imagery), and OrganAMNIST (medical scans). These are combined into a unified dataset with a shared label space of 38 classes. The experts are trained with domain-skewed sampling, where each expert is biased 70% toward a target domain to create specialization.

The heterogeneous experts include:

  • ResNet-50 (CNN) - ResNets excel at extracting precise local features and are highly parameter-efficient compared to ViTs and VLMs. This makes them particularly effective for domains like medical imagery

  • DINOv2 and DINOv3 (SSL) - These models are particularly well-suited for complex visual domains, such as facial recognition

  • CLIP ViT-L/14 (VLM) - VLMs are well-suited for broad visual domains, where visual characteristics correspond closely to human-interpretable concepts

The MLLM router uses Gemma-4-12B, which selects one of five domains, routing the image to the corresponding expert for classification. The system prompt defines five domain aliases: GENERAL (CIFAR10), MEDICAL (OrganAMNIST), GEOGRAPHIC (EuroSAT), FACIAL (FER2013), and UNSURE (a fallback that routes to the best overall model). The router also receives image-quality statistics (blur, brightness, contrast, noise) and uses chain-of-thought reasoning.

Key results show that ARMDIL achieves an impressive 90.78% accuracy on the unified test set, outperforming the strongest individual backbone (the balanced DINOv3 expert) by 1.17%. It also surpasses the majority vote ensemble by 0.31% in overall accuracy and 1.28% in macro F1 score. Compared to a neural network router, the gap in overall accuracy is a mere 0.40%, but ARMDIL achieves this without the training required by a NN-Router.

The paper highlights two major advantages: adaptability and interpretability. For adaptability, Integrating new domains, expert models, or prior knowledge requires only a straightforward modification to the routing prompt. For interpretability, ARMDIL's routing mechanism offers enhanced interpretability by generating explicit, natural-language explanations for its decisions, as demonstrated in Figure 4 where the MLLM provides a step-by-step reasoning trace.

Ablation studies show that self-consistency and image-quality statistics have mixed effects. While self-consistency improves EuroSAT routing accuracy by 3.97%, the overall classification accuracy of the ARMDIL baseline remains superior due to its built-in fault tolerance.

The paper concludes: "ARMDIL effectively navigates the strengths and vulnerabilities of each expert within a heterogeneous ensemble of CNNs, SSL models, and VLMs. Our proposed model outperforms the standard majority vote ensemble and is highly competitive with a specialized neural network-based router, all without requiring any routing-specific training."

Improvements for AI systems

Improvements to AI systems:

  1. Dynamic, interpretable model routing without retraining – Instead of a fixed single model or a trained router, an AI system can use a multimodal LLM agent to select the best-suited vision backbone per input, based on natural-language reasoning about the image’s domain and quality. This eliminates the need for router-specific training and allows instant adaptation to new domains by editing the prompt.

  2. Heterogeneous expert ensembles with domain-skewed specialization – An AI system can combine CNNs (for local/medical features), SSL models (for complex/facial patterns), and VLMs (for broad/human-interpretable concepts), each pre-biased toward a specific domain via 70/30 sampling. This creates complementary specialists that outperform any single backbone.

  3. Unified label space across disparate datasets – The system can merge multiple classification tasks (e.g., objects, faces, satellite, medical) into one shared 38-class space, enabling a single ensemble to handle cross-domain inputs without task-specific fine-tuning.

  4. LLM-based chain-of-thought routing with image-quality metadata – The router can incorporate low-level statistics (blur, brightness, contrast, noise) alongside visual content, and use step-by-step reasoning to decide between GENERAL, MEDICAL, GEOGRAPHIC, FACIAL, or UNSURE. This improves interpretability and allows fault-tolerant fallback to the best overall model when uncertain.

  5. Self-consistency for domain-specific routing refinement – The system can optionally run multiple reasoning passes (self-consistency) to improve routing accuracy for difficult domains like satellite imagery, while keeping the default single-pass mode for speed and overall accuracy.

What the improved AI system can do:

  • Classify images from mixed, unseen distributions (e.g., a photo of a face, a medical scan, a satellite image, and a random object in one batch) with higher accuracy than any single model, and with explicit natural-language explanations for why each image was routed to a particular expert.

  • Add a new domain (e.g., wildlife, industrial defects) by simply adding a new expert and updating the routing prompt—no retraining of the router or ensemble.

  • Operate in low-resource settings by using a lightweight MLLM router (e.g., 12B parameters) and skipping self-consistency when latency matters, while still outperforming majority-vote ensembles.

  • Provide human-auditable decision traces, enabling users to verify or correct routing logic, which is critical for medical, legal, or safety-sensitive applications.

  • Achieve near-neural-network-router performance (within 0.4% accuracy) without any training overhead, making it deployable in zero-shot or few-shot scenarios.

Abstract

Modern image classification models excel when trained on single task-specific datasets but often struggle to generalize across domains and difficulty levels. We propose ARMDIL, an Adaptive Router for Multi-Domain Image classification with LLMs. ARMDIL is an ensemble that uses a multimodal large language model (MLLM) agent to dynamically route each image to the most suitable vision backbone. Our diverse ensemble employs convolutional neural networks (ResNets), self-supervised representation learners (SSL), and vision-language models (VLMs), each trained on a unified label space constructed from multiple image datasets with differing distributions and characteristics. Empirical evaluations illuminate the distinct capabilities and vulnerabilities of each architecture across disparate visual domains. Crucially, we show that ARMDIL effectively navigates these trade-offs, performing competitively with specialized training-based routers. Furthermore, it drastically improves adaptability by allowing new information to be integrated via simple prompt modifications, while enhancing interpretability through natural language reasoning traces. These advances in cross-dataset image classification pave the way for more reliable general-purpose vision systems such as AI assistants and autonomous robots.

Sources

Related papers