Vision Transformers versus convolutional neural networks for fine-grained orchid genus identification in a species-rich, data-poor flora: a controlled benchmark on the Orchidaceae of New Guinea
cs.CV, cs.LG, q-bio.QM
Submitted: 2026-09-21
Updated: 2026-09-21
Comments: 24 pages, 7 figures, 4 tables
Project page: https://birdsheadorchid.id
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
The gist: New Guinea is the world's richest island flora (2,856 orchid species), yet most species are represented by only a handful of photographs, far fewer than direct species-level classification requires.
Terminology
Abstract
New Guinea is the world's richest island flora (2,856 orchid species), yet most species are represented by only a handful of photographs, far fewer than direct species-level classification requires. Methods for fine-grained identification in such species-rich, data-poor floras are needed, and it remains unclear which backbone architecture and pretraining strategy best support them. We built a two-stage system that first predicts the genus of a query photograph, then retrieves visually similar reference images of candidate species using FAISS. We compared four pretrained backbones -- two Vision Transformers (ViTs; DINOv2, BioCLIP 2) and two CNNs (ConvNeXt V2-L, EfficientNetV2-L) -- fine-tuned under an identical protocol on a fixed, species-stratified partition of 16,701 photographs spanning 120 genera and 1,350 species, assessing accuracy, calibration, error structure, species retrieval, and open-set detection of novel genera. DINOv2 attained the best genus performance (macro top-1 66.9%, 95% CI 63.7-70.6; global top-1 88.9%); both ViTs outranked both CNNs, and general-purpose self-supervised pretraining (DINOv2) outperformed domain-matched biological pretraining (BioCLIP 2) by 7.1 points of macro top-1. Errors concentrated on two abundant genera acting as error attractors. DINOv2 embeddings achieved species Recall@5 of 86.6% and genus Recall@5 of 98.7%; temperature scaling reduced every backbone's Expected Calibration Error to about 0.03; and a distance-based open-set gate flagged unseen genera (mean AUROC 0.958). A self-supervised Vision-Transformer backbone combined with embedding retrieval is an effective, deployable strategy for fine-grained identification in species-rich, data-poor floras. The system is released as an open web application (the New Guinea Orchid Identifier), offering a practical template for other hyperdiverse, under-documented taxa.
Sources
- Emerging Properties in Self-Supervised Vision Transformers
- Open-Insect: Benchmarking Open-Set Recognition of Novel Species in Biodiversity Monitoring
- Class-Balanced Loss Based on Effective Number of Samples
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- How Well Do Self-Supervised Models Transfer?
- BioCLIP 2: Emergent Properties from Scaling Hierarchical Contrastive Learning
- On Calibration of Modern Neural Networks
- A Baseline for Detecting Misclassified and Out-of-Distribution Examples in Neural Networks
- Billion-scale similarity search with GPUs
- Revisiting the Calibration of Modern Neural Networks
- DINOv2: Learning Robust Visual Features without Supervision
- An Introduction to Convolutional Neural Networks
- Learning Transferable Visual Models From Natural Language Supervision
- Do Vision Transformers See Like Convolutional Neural Networks?
- BioCLIP: A Vision Foundation Model for the Tree of Life
- EfficientNetV2: Smaller Models and Faster Training
- The iNaturalist Species Classification and Detection Dataset
- Open-Set Recognition: a Good Closed-Set Classifier is All You Need?
- ConvNeXt V2: Co-designing and Scaling ConvNets with Masked Autoencoders
- Constructing Balance from Imbalance for Long-tailed Image Recognition
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models