Automated binary classification of hazelnut X-ray images: A deep-learning benchmark for quality assessment
Giancarlo Sportelli, Nicola Belcari, Roberta Pace, Umberto Bernardo, Sharmin Sultana, Alessandra Toncelli, Matteo Giaccone
University of Pisa · National Institute for Nuclear Physics · National Research Council of Italy · National Research Council of Italy
cs.CV, cs.LG, physics.app-ph
Submitted: 2026-08-12
Updated: 2026-08-13
Comments: 26 pages (including 5 pages of supplementary material), 4 figures, 5 tables. Dataset available at https://doi.org/10.5281/zenodo.21739932
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 100/100
The gist: This study presents a benchmark for binary hazelnut quality classification (healthy versus defective) using deep learning on X-ray images.
Terminology
Summary
This study presents a benchmark for binary hazelnut quality classification (healthy versus defective) using deep learning on X-ray images. The dataset consisted of 799 segmented single-kernel X-ray images (224 × 224 pixels, grayscale) of Corylus avellana var. pontica, cv. Anakliuri, grouped into 101 acquisition units. The samples were harvested in 2024 from an organic orchard in Zugdidi, Georgia, and dried to 6% moisture content. Quality was initially assessed through external visual inspection following UNECE standards, with classes including healthy (H), stink bug-damaged (C), rotten (R), and oil-rancidity (O). After X-ray imaging, externally healthy kernels were manually opened to identify hidden rot (HR). Under the initial annotation condition, the dataset contained 138 healthy (17.3%) and 661 defective (82.7%) kernels. Following expert reassessment of 15 kernels that had appeared externally healthy but had been assigned to defect classes after internal inspection, the binary labels of these 15 kernels were changed from defective to healthy, resulting in 153 healthy (19.1%) and 646 defective (80.9%) kernels under the reassessed annotation condition.
Seven single-model configurations were evaluated: BinaryNutCNN (a lightweight custom CNN with 0.42 million parameters) trained with three loss configurations (binary cross-entropy, focal loss with alpha=0.75 and gamma=1.0, and BCE with pos weight=2.0), Swin Transformer Tiny (both full fine-tuning and frozen-backbone variants), EfficientNet-B0, and ResNet-18. Ten ensembles were constructed by combining a custom-CNN configuration with a pretrained backbone using probability aggregation (average or maximum) at inference. The evaluation used a group-wise split-rotation protocol with five data splits generated using different random seeds (42, 123, 456, 789, 2024), with a 70%/15%/15% train/validation/test partition stratified by majority class per acquisition unit. The model initialization seed was held constant across runs. Decision thresholds were selected on the validation set via grid search from 0.10 to 0.90 in steps of 0.02, and performance was assessed deterministically on validation and test sets. Training used up to 50 epochs with early stopping (patience 15 epochs), a WeightedRandomSampler for class imbalance, and augmentation limited to random horizontal and vertical flips on the training set only.
Under the expert-reassessed annotation condition, the average-probability ensemble of the binary cross-entropy-trained convolutional neural network and frozen Swin Transformer achieved the highest mean balanced accuracy of 86.3% ± 1.8% across the five seeds. This configuration showed one of the lowest levels of cross-split variability, with mean defect recall of 79.4% ± 2.0% and mean healthy recall of 93.1% ± 5.0%. The absolute error counts averaged 20.6 ± 2.0 false negatives per split (range: 18–24) and 1.8 ± 1.2 false positives per split (range: 0–3), with test folds comprising 126–128 kernels. Validation-selected decision thresholds ranged from 0.42 to 0.78 across splits. The remaining top-ranked ensembles achieved mean balanced accuracies ranging from 83.1% to 85.1%. Ensemble configurations occupied the seven highest-ranking positions, although several obtained comparable results. Notably, ensembles incorporating the frozen Swin Transformer remained competitive, indicating that full backbone fine-tuning was not required to achieve strong performance.
Expert reassessment changed mean balanced accuracy by −0.5 to +3.4 percentage points across the 17 methods, with a mean change of +1.4 percentage points. Thirteen of the seventeen methods showed a positive mean change (up to +3.4 pp for ens max focal full and +3.1 pp for ens avg bce frozen), while four methods showed small negative changes within one standard deviation of zero (−0.1 to −0.5 pp). At the individual-split level the effect was not uniform, with some seeds performing slightly worse under the reassessed condition. Cross-split variability was essentially unchanged between conditions (average standard deviation 3.44% under the initial condition and 3.46% under the reassessed condition), with 9 methods showing a marginal increase and 8 a marginal decrease.
Split-to-split variability was substantial. For the top-ranked method, ens avg bce frozen, performance ranged from 83.8% to 88.9%, corresponding to a span of 5.1 percentage points, while other methods exhibited wider ranges. Methods tended to improve or decline on the same splits, indicating that evaluation performance is strongly influenced by test-fold composition, particularly by which acquisition units contribute healthy samples. No method achieved the highest balanced accuracy on every split, and split-to-split variability was of the same order as the small differences in mean balanced accuracy among leading configurations. The ordering in Table 4 should therefore be interpreted as a ranking based on five-split means rather than evidence of consistent per-split dominance.
As a supplementary sensitivity analysis, the complete benchmark was repeated using nut-level splitting, which allowed kernels from the same acquisition unit to be distributed across training, validation, and test subsets. Across the 17 evaluated methods, mean balanced accuracy was 81.5% under nut-level splitting and 82.8% under group-wise splitting, providing no indication that acquisition-specific characteristics produced an optimistic performance advantage. The group-wise protocol was retained as the primary evaluation strategy because it prevents acquisition-level information leakage and provides a more conservative estimate of generalization to previously unseen acquisitions.
The study concludes that two-dimensional X-ray imaging combined with deep-learning ensembles can support the non-destructive separation of healthy and defective hazelnut kernels. Its main contribution lies not in identifying a universally superior architecture, but in establishing a rigorous evaluation framework in which annotation quality and independence among acquisition groups are integral to model assessment. The results highlight both the potential of deep learning for automated X-ray-based hazelnut quality assessment and the importance of rigorous evaluation and label curation in small, imbalanced agricultural imaging datasets. The authors note that the reported performance is specific to the experimental X-ray system used and should not be regarded as scanner-independent estimates, and that industrial translation will require a dedicated, compact, and high-throughput platform with prospective validation on independent industrial lots.
Improvements for AI systems
Improvements to AI Systems:
-
Ensemble probability aggregation (average vs. max) with heterogeneous backbones – Combine a lightweight custom CNN with a frozen transformer (e.g., Swin-Tiny) using average probability fusion. This yields higher balanced accuracy (86.3%) and lower variance than any single model, while avoiding costly full fine-tuning.
-
Frozen-backbone feature extraction for small, imbalanced datasets – Use a pretrained transformer with its weights frozen, only training a classification head. This reduces overfitting and achieves competitive performance without the need for large-scale fine-tuning, which is critical when training samples are scarce (799 images).
-
Class-imbalance-aware loss functions with tunable hyperparameters – Implement focal loss (alpha=0.75, gamma=1.0) or BCE with pos weight=2.0, combined with a WeightedRandomSampler. These mitigate the severe class imbalance (80.9% defective vs. 19.1% healthy) and improve recall for the minority healthy class.
-
Group-wise split-rotation with multiple seeds for robust evaluation – Replace random splits with acquisition-unit-level grouping to prevent data leakage. Use 5 different seeds and report mean ± std across splits, enabling reliable model selection and uncertainty quantification.
-
Validation-based decision threshold optimization via grid search – Automatically select the optimal classification threshold (range 0.10–0.90, step 0.02) on the validation set per split. This adapts to varying class distributions and improves balanced accuracy without manual tuning.
-
Label curation via expert reassessment – Integrate a feedback loop where model predictions on externally healthy samples are verified by internal inspection (e.g., opening kernels to detect hidden rot). This corrects mislabeled training data, improving mean balanced accuracy by +1.4 percentage points across methods.
-
Early stopping with patience and deterministic evaluation – Use up to 50 epochs with patience 15, and fix the model initialization seed. This ensures reproducible results and prevents overfitting, especially important for small datasets.
-
Augmentation limited to flips on training only – Apply only random horizontal and vertical flips to avoid distorting X-ray features (e.g., density gradients). This preserves physical realism while providing regularization.
What the Improved AI System Can Do:
-
Non-destructively classify hazelnut kernels as healthy or defective from 2D X-ray images with 86% balanced accuracy, detecting hidden internal defects (e.g., rot, stink bug damage, rancidity) that are invisible externally.
-
Operate reliably on small, imbalanced datasets (e.g., <1000 images) by using frozen backbones, ensemble averaging, and class-weighted losses, reducing the need for large annotated corpora.
-
Generalize to new acquisition batches (e.g., different harvests or orchards) by using group-wise splitting that prevents leakage from the same source unit.
-
Provide uncertainty-aware predictions by reporting per-split variance and threshold ranges, allowing operators to set sensitivity based on cost of false negatives vs. false positives.
-
Automate quality control in food processing – sort defective kernels before packaging, reducing waste and ensuring product safety, with a path toward high-throughput industrial X-ray platforms.
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models