Retrieval-Augmented Vision Foundation Models for Robust Leukemia Cell Classification across Multiple Microscopy Datasets
Carlos Zamora, Hiram Zuniga, Ulises Orozco-Rosas, Kenia Picos
CETYS University
eess.IV, cs.CV, cs.LG
Submitted: 2026-08-11
Updated: 2026-08-12
Comments: Accepted at SPIE Optics + Photonics 2026 for oral presentation. 23 pages, 12 figures, 9 tables
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 75/100
The gist: This work presents a robust framework for leukemia classification across multiple heterogeneous datasets using a two-stage pipeline with a pretrained vision foundation model.
Terminology
Summary
This work presents a robust framework for leukemia classification across multiple heterogeneous datasets using a two-stage pipeline with a pretrained vision foundation model. Stage 1 performs binary classification (leukemia vs. non-leukemia) and is trained using 122,167 single-cell images. Stage 2 is conditionally applied to Stage 1 positives to perform subtype classification into Acute Lymphoblastic Leukemia (ALL) and Acute Myeloid Leukemia (AML), trained using 69,400 single-cell images. Labels are harmonized across five heterogeneous datasets to enable cross-dataset training, and performance is evaluated on a held-out dataset protocol to assess domain-shift generalization. Within this pipeline, three encoders are benchmarked (DinoBloom, pretrained on single-cell images; BiomedCLIP, pretrained on biomedical data; and CLIP as a general-purpose model) under linear probing, Low-Rank Adaptation (LoRA), and a Retrieval-Augmented Classification (RAC) module that retrieves the top-k most similar cell images to provide cytomorphological grounding. The objective is to quantify how much domain-specific pretraining contributes to performance under domain shift, and whether cost-effective adaptation and retrieval can be a viable alternative to expensive domain-specialized pretraining. The held-out protocol additionally serves as a diagnostic tool, revealing when classification performance is attributable to dataset-specific artifacts rather than to cytomorphological features.
The study addresses the following research gap: how much does pretraining affect model performance and robustness under domain shift, and how much can fine-tuning and other techniques make up for this difference? Quantifying this impact allows us to evaluate whether cost-effective methods such as PEFT and retrieval augmentation can be a viable alternative to expensive domain-specialized pretraining for leukemia classification.
The main contributions of this work are described as follows:
• A two-stage classification pipeline, performing a binary classification at Stage 1 on leukemia vs non-leukemia, and then performing a subtype classification (ALL/AML) on leukemia positives for Stage 2, reducing the broad classification task into two steps.
• A systematic benchmark that quantifies the impact of pretraining specialization on a foundation model for leukemia classification, evaluating models pretrained on single-cell images, biomedical data, and general-purpose data within the same two-stage pipeline. The benchmark further isolates the contribution of adaptation strategies (linear probe and LoRA) and RAC, quantifying how much these techniques compensate for differences in pretraining.
• A held-out dataset evaluation protocol to assess performance under an unseen domain, instead of using a test split from domains seen during training. Cross-dataset training is performed using the remaining heterogeneous datasets (122,167 images for Stage 1 and 69,400 for Stage 2), comprising different acquisition protocols, staining, and labeling paradigms. This evaluation protocol also exposes dependencies on the acquisition source that a within-domain test split cannot reveal.
Five publicly available datasets containing cropped single-cell images of white blood cells (WBCs) were unified into a single dataset, consisting of three main classes: ALL, AML, and normal cells. The purpose of this merge was to improve the framework model’s generalization capability by introducing varying image quality, staining protocols, and image acquisition methods, increasing overall sample diversity. After preprocessing, the final dataset contained 122,427 single-cell images. The datasets used were: C-NMC 2019 (D1, 12,528 usable images), AML-Cytomorphology (D2, 18,127 usable images), Peripheral Blood Cell (D3, 10,298 usable images), AML-Cytomorphology MLL Helmholtz (D4, 81,214 usable images), and ALL-IDB2 (D5, 260 usable images).
To mitigate the large domain shift and reduce the risk of the model learning to differentiate the ALL class by memorizing image elements unrelated to cell morphology, each image from C-NMC 2019 was processed. The black background was removed and replaced with a randomly selected area from a healthy peripheral blood smear sample, containing mainly RBCs. To increase generalization, the background was randomly modified through rotations and scaling. Additionally, Reinhard color normalization was applied to improve color consistency between datasets.
Across the five heterogeneous datasets used for training and evaluation, annotation schemes differ substantially. To enable cross-dataset training and held-out dataset evaluation, this work harmonizes these heterogeneous labels into three classes: normal, ALL, and AML. All genetic AML subtypes were collapsed into a single AML class. A uniform inclusion criterion was applied: only non-ambiguous nucleated leukocytes were used, since these are the cells relevant to leukemia diagnosis and the only cell type consistently present across all five datasets. Consequently, this study excluded non-leukocyte samples (platelets and erythroblasts), as they are not leukocytes, and smudge cells (labeled as KSC in D2), as they are ruptured cells that cannot be reliably identified. Immature granulocytes (promyelocytes, myelocytes, and metamyelocytes) were excluded because they are labeled inconsistently across datasets. Some sources label them as normal cells and others as leukemic. For the myeloid leukemic class, positive samples were restricted to blasts (myeloblasts and monoblasts). These rules and criteria were applied uniformly across all five datasets to ensure a consistent preprocessing pipeline before training and evaluation.
In this work, leukemia classification is decomposed into two sequential stages, each implemented as an independent linear classifier on top of a vision encoder: DinoBloom, BiomedCLIP, or CLIP. Stage 1 performs a binary classification (leukemia vs. non-leukemia) over the full harmonized corpus of 122,427 single-cell images, collapsing the AML and ALL classes into a single leukemia class. Stage 2 performs subtype classification (ALL vs. AML), applied conditionally to the samples predicted as leukemic by Stage 1, over the 72,824 leukemic single-cell images that remain after removing samples labeled as normal. Decoupling the task this way allows Stage 1 to leverage every labeled image, including datasets that provide only normal single-cell images, rather than being restricted to sources with subtype annotations, augmenting the corpus from 72,824 images to 122,427 images. The two stages are trained and evaluated as separate classifiers under their own held-out datasets, avoiding contamination or data leakage between stages.
Three vision encoders from different pretraining specializations were trained and compared within the same pipeline. DinoBloom is a DINOv2-based model developed by MarrLab and pretrained on 13 publicly available datasets of single-cell images from peripheral blood and bone marrow, serving as the specialized model. BiomedCLIP is pretrained on biomedical image and text pairs extracted from PubMed Central articles and represents the semi-specialized model, since it is pretrained on a biomedical domain but not as specific as single-cell images. Lastly, CLIP is the general-purpose baseline, pretrained on natural image and text pairs. The versions used for these three vision encoders were: DinoBloom-B, which produces 14×14 image patches, 768-dimensional embeddings, and has 85.7 million parameters. BiomedCLIP produces 16×16 image patches, 512-dimensional embeddings, and has 86.2 million parameters, and CLIP produces 32×32 image patches, 512-dimensional embeddings, and has 87.8 million parameters. The purpose of choosing these base models was to isolate the pretraining domain as the main source of variation between them, since the three encoders are consistent in parameter quantity; no result can be attributed to encoder size. For the two vision-language models (BiomedCLIP and CLIP), only the vision tower is used, since the pipeline requires no text input. All embeddings are L2-normalized before classification.
Two adaptation strategies are evaluated using these three encoders. Linear probing involves freezing the encoder and training only a logistic regression classifier on the precomputed embeddings, using class weighting to compensate for class imbalance. This configuration isolates the representational quality of the pretrained encoder, since no parameter of the backbone is modified, only the head classifier. No augmentation was applied for linear probing, since embeddings were extracted once with deterministic preprocessing, and the encoder never processes the images again during training. In Low-Rank Adaptation, low-rank update matrices are injected into the encoder’s attention projection layers while the pretrained weights remain frozen. The LoRA parameters and a linear classification head are trained jointly. LoRA was configured with the following hyperparameters: rank r = 8, scaling factor 16, and dropout 0.05. This results in 296,450 trainable parameters for DinoBloom, 295,938 for BiomedCLIP, and 443,394 for CLIP, a small fraction of the approximately 86 million parameters of each encoder. Training used the AdamW optimizer with a learning rate of 1×10−4, a batch size of 128, and class-weighted cross-entropy, for a fixed budget of three epochs with mixed-precision computation. Light augmentation consisting of horizontal and vertical flips and mild color jitter was applied during LoRA training. The same LoRA hyperparameters and training budget were applied uniformly across the three encoders and both stages to ensure comparability. The held-out datasets were never used during training or hyperparameter selection.
The RAC module works on the embedding space produced by the frozen or LoRA-adapted encoder. Given a query image, its embedding is compared by cosine similarity against an embedding bank comprising the L2-normalized training embeddings produced by the encoder. The top-k most similar samples make a similarity-weighted vote over classes, making a retrieval distribution pret. This distribution is then fused with the linear classifier distribution pprobe in log space, as defined in Eq. (2). Where α weights the contribution of the retrieval branch. In the linear probe configurations, the α value was selected on an internal validation split derived from the training datasets, searching over the grid 0, 0.25, 0.5, 0.75, 1. In contrast, for the LoRA configurations, α was fixed at 0.5 due to computational budget. The held-out datasets were never used for tuning. The predicted class corresponds to the highest resulting logit, and the k value was set to 20 in all experiments. This RAC module is implemented as an exact search in PyTorch. Since embeddings are L2-normalized, cosine similarity reduces to a dot product between the query and the embedding bank, and the top-k neighbors are obtained directly. No approximate nearest neighbor index or external retrieval library such as FAISS was used.
To assess model performance under domain shift, each stage is evaluated on a held-out dataset excluded entirely from training, rather than on a test split derived from domains already seen by the model. Each stage uses its own held-out dataset, since the two stages are trained and evaluated as independent classifiers. Stage 1 is evaluated on the ALL-IDB2 dataset, which contains both leukemic and healthy cells in a balanced proportion (130 and 130 single-cell images). This dataset is not part of the reported pretraining corpus of DinoBloom, which strengthens the evaluation, since it represents a domain unseen by the domain-specific encoder. Stage 1 training used the four remaining datasets, consisting of 122,167 single-cell images. Stage 2 is evaluated on the combination of ALL-IDB2 (representing the ALL class) and AML-Cytomorphology (representing the AML class), because, to the best of the authors’ knowledge, no single public dataset contains both subtypes. This results in a test set of 3,424 images (130 ALL and 3,294 AML). Stage 2 training was performed on the remaining two datasets, consisting of 69,400 leukemic single-cell images. In both stages, the held-out datasets are excluded from training to avoid contamination or data leakage. Neither classifier is reused across stages, avoiding data leakage between them. As a control experiment, Stage 2 is additionally evaluated under a conventional random 80/20 split stratified by class, using the same encoders and linear classifier. This within-domain split quantifies how much of the reported performance depends on dataset-specific artifacts rather than cell morphology patterns. Performance is reported as accuracy, macro-averaged recall, and macro-averaged F1-score.
In Stage 1, under held-out dataset evaluation, pretraining specialization accounts for a 0.2577 accuracy gap between DinoBloom (domain-specific) and CLIP (general-purpose) when both are evaluated with a frozen encoder (linear probing). This gap narrows to 0.0193 once LoRA adaptation is applied. Additionally, CLIP adapted with LoRA (0.9115) surpasses DinoBloom without adaptation (0.8923). The best overall result (0.9423) was achieved by combining the domain-specific encoder with retrieval augmentation and linear probing, without any fine-tuning. Cost-effective adaptation methods can compensate for most of the advantage provided by expensive domain-specific pretraining in this leukemia and subtyping classification task, although the highest performance was still achieved by combining both specialized pretraining with RAC. Therefore, cost-effective adaptation benefits not only general-purpose encoders but also specialized ones, and is a viable alternative if expensive pretraining is not feasible due to time or computational resources.
The held-out dataset evaluation made visible a limitation that a test split within the domain would not expose in this experimental setting. In Stage 2, subtype classification collapses to the majority class under held-out dataset evaluation, while the same models and data achieve near-perfect accuracy using a random within-domain 80/20 split. This is attributable to each subtype class coming from a single dataset or imaging source during training, which allows the encoder to rely on dataset-specific artifacts instead of cytomorphological features. Therefore, future work should focus on adding more data sources to force the encoder to learn cellular patterns rather than shortcuts based on smear backgrounds. Another direction would be to normalize the backgrounds in all single-cell images across datasets, which could be done by applying the same synthetic background. This would force the encoder to learn cellular patterns instead of dataset-specific artifacts, providing a more consistent comparison.
Improvements for AI systems
Improvements to AI Systems:
- Two-Stage Hierarchical Classification Pipeline
-
Implement a sequential classifier that first performs binary screening (disease vs. healthy) using all available data, then conditionally applies subtype classification only to positive cases.
-
This reduces class imbalance, leverages larger datasets for the first stage, and improves computational efficiency by avoiding subtype classification on negatives.
- Domain-Shift Robustness via Cross-Dataset Training and Held-Out Evaluation
-
Train on multiple heterogeneous datasets with harmonized labels (e.g., unifying staining protocols, acquisition methods, and label schemes) to force the model to learn morphology rather than dataset-specific artifacts.
-
Evaluate on a completely unseen dataset (not a random split from training domains) to measure true generalization and detect shortcut learning.
- Pretraining Specialization-Aware Model Selection
-
Use domain-specific pretrained encoders (e.g., DinoBloom for single-cell images) when available, but quantify the performance gap against general-purpose models (e.g., CLIP) to decide if expensive pretraining is justified.
-
For resource-constrained settings, apply Low-Rank Adaptation (LoRA) to general-purpose encoders to recover most of the performance gap (e.g., 0.2577 accuracy gap reduced to 0.0193) without full fine-tuning.
- Retrieval-Augmented Classification (RAC) for Cytomorphological Grounding
-
Integrate a k-nearest-neighbor retrieval module (top-k=20) that compares query embeddings against a training embedding bank, fusing the retrieval distribution with the classifier output via log-space weighting (α).
-
This provides interpretable grounding by referencing similar cell images, improving accuracy (e.g., best result 0.9423 with DinoBloom + RAC) and reducing reliance on dataset-specific backgrounds.
- Adaptive Fusion of Classifier and Retrieval Signals
-
Dynamically weight the contribution of the linear classifier vs. retrieval branch (α tuned on validation, or fixed at 0.5 for LoRA) to balance memorization of training patterns with similarity-based reasoning.
-
This allows the system to leverage both global decision boundaries and local exemplar evidence.
- Background Normalization and Synthetic Augmentation
-
Replace non-informative backgrounds (e.g., black regions in C-NMC 2019) with randomly sampled healthy blood smear backgrounds, plus rotations/scaling, to prevent the model from memorizing background artifacts.
-
Apply Reinhard color normalization across datasets to reduce staining variability, improving cross-dataset consistency.
- Label Harmonization and Inclusion Criteria
-
Collapse inconsistent subtype labels (e.g., genetic AML subtypes) into unified classes, and exclude ambiguous cells (e.g., smudge cells, immature granulocytes) to ensure clean training signal.
-
Restrict positive samples to clearly defined blasts (myeloblasts/monoblasts) for myeloid leukemia, reducing label noise.
- Diagnostic Evaluation Protocol for Shortcut Detection
-
Compare held-out dataset performance against within-domain random splits to quantify reliance on dataset-specific artifacts (e.g., Stage 2 collapsed to majority class on held-out data but achieved near-perfect accuracy on random splits).
-
Use this discrepancy to flag when models are exploiting shortcuts, guiding future data collection or normalization efforts.
What the Improved AI System Can Do:
-
Accurately classify leukemia vs. healthy cells and subtype (ALL vs. AML) across unseen clinical domains, even when training data come from different scanners, stains, and institutions.
-
Achieve high performance with limited computational resources by using LoRA-adapted general-purpose encoders, making deployment feasible in low-resource labs.
-
Provide explainable decisions by retrieving and displaying the most similar training cell images that influenced the classification.
-
Detect and avoid shortcut learning by identifying when performance drops on held-out datasets, prompting corrective actions like background normalization or additional data collection.
-
Scale to new datasets with minimal retraining by harmonizing labels and using retrieval augmentation to adapt to novel imaging conditions.
Sources
- Imbalanced Domain Generalization for Robust Single Cell Classification in Hematological Cytomorphology
- Parameter-Efficient Fine-Tuning for Large Models: A Comprehensive Survey
- LoRA: Low-Rank Adaptation of Large Language Models
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
- Retrieval Augmented Classification for Long-Tail Visual Recognition
Related papers
- Revisiting Integration of Image and Metadata for DICOM Series Classification: Cross-Attention and Dictionary Learning
- VesselSDF: Distance Field Priors for Vascular Network Reconstruction
- cSVR: Convolutional Slice-to-Volume Reconstruction
- NAIMA: Semantics Aware RGB Guided Depth Super-Resolution
- AneumoBench: A Source-Linked Benchmark for Synthetic-Geometry Transfer in Aneurysm CFD
- RETO: A Rotary-Enhanced Transformer Operator for High-Fidelity Prediction of Automotive Aerodynamics