A comparison of CNN architectures for Alzheimer's disease detection in single-view MRI scans

arXiv:2608.11762 · eess.IV, cs.CV, cs.LG · Submitted 2026-08-12 · Read on arXiv

Hiram Zuniga, Ulises Orozco-Rosas, Kenia Picos

CETYS University

eess.IV, cs.CV, cs.LG

Submitted: 2026-08-12

Updated: 2026-08-13

Comments: Accepted at SPIE Optics + Photonics 2026 for oral presentation. 14 pages, 7 figures, 8 tables

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 75/100

The gist: This paper proposes a benchmark that evaluates ten different convolutional neural network (CNN) architectures (including ResNet, DenseNet, MobileNet, EfficientNet, and VGG family models) under the

Terminology

Summary

This paper proposes a benchmark that evaluates ten different convolutional neural network (CNN) architectures (including ResNet, DenseNet, MobileNet, EfficientNet, and VGG family models) under the same held-out test split protocol. A two-stage transfer learning and full fine-tuning pipeline is introduced to perform training using a class-balanced subset (3,900 images) derived from the OASIS medical imaging dataset, comprising 86,437 single-view MRI brain scans labeled into four classifications of Alzheimer’s disease: Non-Demented, Very Mild Dementia, Mild Dementia, and Moderate Dementia. The best results were achieved by VGG16, with a 0.9637 validation accuracy and a 0.9533 test accuracy score. A key finding documented in this work is the difficulty of classifying the transition from Non-Demented to Very Mild Demented stages, observed consistently across all ten architectures.

Improvements for AI systems

Improvements to AI Systems:

  1. Class-Imbalance-Aware Training Protocol: Integrate the paper’s two-stage transfer learning pipeline (feature extraction + full fine-tuning) with class-balanced sampling. This improves the system’s ability to learn from rare or underrepresented disease stages (e.g., Moderate Dementia) without overfitting, leading to more stable performance across all severity levels.

  2. Transition-Stage Sensitivity Detection: Add a dedicated loss function or auxiliary classifier that penalizes misclassification between adjacent disease stages (e.g., Non-Demented vs. Very Mild). This makes the AI system more conservative and clinically useful by reducing false negatives for early-stage Alzheimer’s, which is critical for timely intervention.

  3. Architecture-Agnostic Benchmarking Module: Build an automated model-selection layer that runs the same held-out test split and fine-tuning pipeline across multiple CNN families (ResNet, EfficientNet, etc.). This allows the AI system to dynamically choose the best-performing backbone for a given medical imaging task, rather than relying on a single pre-chosen architecture.

  4. Confidence Calibration for Early-Stage Diagnosis: Use the paper’s finding that all ten models struggle with the Non-Demented → Very Mild transition to implement uncertainty quantification (e.g., Monte Carlo dropout or ensemble variance). The improved system will output a confidence score alongside the prediction, flagging low-confidence cases for human review, thereby reducing diagnostic errors in ambiguous early-stage scans.

What the Improved AI System Can Do:

  • Achieve higher and more consistent accuracy across all four Alzheimer’s stages, especially improving recall for Very Mild Dementia (the most clinically challenging category).

  • Provide a transparent, reproducible benchmark for any new CNN architecture, enabling rapid comparison and deployment of the best model for a given dataset.

  • Generate calibrated risk scores that alert clinicians when a scan is borderline between healthy and early disease, supporting proactive monitoring rather than reactive diagnosis.

  • Automatically adapt its training strategy (e.g., rebalancing or stage-weighted loss) when new data arrives, maintaining robustness without manual retuning.

Sources

Related papers