Domain-Aware Lightweight Spectral-Grouped Convolutions for Hyperspectral Fish Freshness Classification
Kazi Nabiul Alam, Pooneh Bagheri Zadeh, Akbar Sheikh-Akbari
Leeds Beckett University
eess.IV, cs.AI, cs.CV, eess.SP
Submitted: 2026-08-12
Updated: 2026-08-13
Comments: Accepted at BMVC'2026
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 75/100
The gist: SGNet (Spectral-Grouped Network) is a lightweight architecture for hyperspectral fish freshness classification that separates spectral and spatial feature extraction using grouped convolutions and a
Terminology
Summary
SGNet (Spectral-Grouped Network) is a lightweight architecture for hyperspectral fish freshness classification that separates spectral and spatial feature extraction using grouped convolutions and a depthwise spatial pathway. A dual attention mechanism that couples channel-wise squeeze-and-excitation with spatial gating adaptively highlights informative features. SGNet achieves 97.8% classification accuracy and 0.64 days mean absolute error (MAE) with just 4.75M parameters when tested on our newly developed 16-day refrigerator-stored salmon fillet dataset. Ablation studies validate the contribution of each component, while comparisons demonstrate a five- to eighteen-fold parameter reduction relative to ResNet-50 and Vision Transformers. Our findings indicate that domain-aware design supports precise, real-time freshness prediction for industrial implementation.
The study introduces SGNet, a framework designed around the structure of hyperspectral food data rather than adapted from a remote sensing or natural image backbone. Rather than proposing new primitives, our contribution is a principled composition of efficient operators that jointly respects spectral dominance, ordinal label structure, and severe sample scarcity, together with the analysis and dataset needed to validate it in this regime.
The contributions are: an explicit spectral-spatial factorisation in which grouped pointwise convolution confines cross-band interaction to local channel subsets while a parallel depthwise pathway captures spatial structure, preventing premature entanglement of unrelated wavelength responses; a lightweight dual attention mechanism coupling channel-wise excitation with spatial gating at below 5% of the block parameters, giving selective emphasis without the data appetite of self-attention; a hierarchical encoder that attains 97.8% accuracy with only 4.75M parameters and is benchmarked against widely used convolutional and transformer architectures retrained under identical conditions; and a curated 16-day refrigerated salmon dataset with strict pack-level separation, addressing the scarcity of leakage-controlled hyperspectral freshness data.
The hyperspectral fish freshness classification task presents three principal obstacles. First, spectral dominance means that freshness information is carried primarily by wavelength-dependent reflectance patterns rather than by spatial texture. Second, temporal smoothness requires that representations of neighbouring days remain close in feature space, so that an error of one day is treated as far less severe than an error of several days. Third, the available data are limited in number and are subject to additional sources of variation and noise. Generic CNNs that combine spectral and spatial information without regard for these constraints perform poorly, because they entangle correlated bands prematurely and overfit to spurious patterns.
The proposed design therefore contains two independent processing paths that use grouped convolutions together with lightweight attention to suppress unimportant variation while performing spectral and spatial interaction in a controlled manner.
The core computational unit is the Spectral–Grouped CNN (SGNet) block, which processes spectral and spatial information through dedicated pathways. Cross-channel interaction is performed via grouped pointwise convolution with g = 8 groups, GELU activation, and the output is further processed by a two-layer MLP with expansion ratio 2. Grouped convolution constrains cross-band interaction to local channel subsets, preventing premature entanglement of unrelated spectral features and reducing sensitivity to noise. Local spatial context is aggregated independently per channel using depthwise separable convolution with kernel size k = 7, capturing local tissue structure and spatial patterns while preserving spectral independence. A lightweight dual attention mechanism inspired by squeeze-and-excitation and convolutional attention computes channel-wise attention via squeeze-and-excitation with bottleneck reduction ratio r = 4, and spatial attention via a 1×1 convolution producing a spatial attention map. The gated output combines both attention mechanisms with element-wise multiplication. Unlike self-attention, this mechanism introduces minimal computational overhead (less than 5% of block parameters) and remains stable under limited data regimes. The final output is obtained via a residual feed-forward transformation with two pointwise convolutions with expansion ratio 4 and GELU activation.
The complete network stacks SGNet blocks across three stages with progressively increasing channel dimensions 64, 128, 256 and correspondingly decreasing spatial resolutions 562, 282, 142. Block depths are set to 3, 3, 4 for stages 1–3 respectively. Downsampling between stages is performed via layer-normalised strided convolution. This hierarchical design enables the model to capture freshness-related cues across multiple spatial scales, ranging from fine-grained local tissue variation to coarse global structural change indicative of degradation.
The model is trained with standard cross-entropy loss, which we adopt as a deliberately simple and strong baseline; this isolates the effect of the architecture from that of the loss. We evaluate using ordinal-aware metrics including mean absolute error (MAE), accuracy (Acc), and quadratic weighted kappa (QWK), which better reflect the continuous nature of freshness evolution and penalise predictions in proportion to their distance from the ground truth.
The curated salmon hyperspectral dataset consists of N = 800 HSI cubes acquired over a 16-day refrigerated storage period (K = 16 classes), where day 6 is the labelled expiry date. Each cube comprises C = 462 spectral bands spanning the visible and near-infrared range. The dataset is split into training (560), validation (112), and test (128) sets with balanced class distributions. To avoid data leakage, distinct fish packs were assigned to each split (35 packs for training, 7 for validation, and 8 for testing), so that no cube from a given pack appears in more than one split.
The model achieves 97.8% classification accuracy with an MAE of only 0.64 days, which indicates that predictions are either correct or deviate by at most one storage day on average. The distribution of error across storage days shows that misclassifications are concentrated in the late spoilage window (days 14–16), which is more than eight days after the labelled expiration date (day 6) and is consistent with the progressive biological degradation of the tissue over time. Days within the commercially relevant pre- and peri-expiry window are classified essentially without error.
SGNet attains the highest accuracy and the lowest ordinal error while using the smallest parameter budget, and the margin over the larger models is most pronounced for the transformer baselines. This pattern is consistent with the well-documented data appetite of attention-based models: in the small-sample, high-dimensional regime that characterises hyperspectral food data, the inductive biases of grouped and depthwise convolution are more valuable than raw capacity. The comparison therefore isolates the benefit of domain-aware factorisation rather than of scale.
Days 1–14 achieve near-perfect classification (F1 ≥ 96%), while days 15–16 show marginally reduced performance (F1 between 87% and 91%) owing to overlapping degradation markers in the advanced spoilage stages.
Ablation studies confirm that every component contributes positively, and removing the depthwise spatial path or the dual attention degrades both accuracy and ordinal error more than removing grouped convolution. Grouped convolution with g = 8 provides the most favourable spectral-efficiency tradeoff, since the ungrouped variant (g = 1) both increases the parameter count and lowers accuracy. The dual attention block improves accuracy at negligible parameter cost. Scaling the network down to channel widths 48, 96, 192 reduces capacity at a modest cost in accuracy, whereas the deeper 4, 4, 6 variant improves MAE slightly but inflates the parameter budget, indicating that the chosen configuration sits at a sensible operating point on the accuracy–efficiency curve.
SGNet achieves a 5.4× parameter reduction relative to ResNet-50 and an 18× reduction relative to ViT-B while maintaining the highest throughput, which makes it well suited to real-time industrial deployment on resource-constrained hardware. Throughput scales favourably with batch size and remains comfortably within real-time requirements for an inspection line, while single-cube latency is low enough for interactive use at the point of acquisition.
The experiments support the central claim that domain-aware factorisation, rather than greater capacity, is the more productive route to accurate hyperspectral freshness assessment. The concentration of residual error in the late spoilage stages (days 15–16) is not a failure mode so much as a reflection of biology, since degradation markers in advanced spoilage overlap and the spectral separation between adjacent late days is genuinely small. The gap between SGNet and the transformer baselines widens precisely in the small-sample regime, which indicates that the inductive biases encoded by grouped and depthwise convolutions are valuable when data are scarce. The dual attention mechanism delivers a measurable accuracy gain for a negligible parameter cost, which suggests that selective emphasis, rather than dense global mixing, is sufficient for this task.
Several limitations remain. The dataset, although carefully pack-separated to prevent leakage, is drawn from a single species under controlled refrigerated storage, so generalisation across species, storage conditions, and acquisition instruments is not yet established. The current model is trained with cross-entropy and evaluated with ordinal metrics; incorporating an explicit rank-consistent objective such as CORAL or its successor CORN into training is a natural extension that may further reduce large-distance errors. Cross-species transfer, explicit ordinal training, and an extension to additional food matrices are identified as the most promising directions for future work.
Improvements for AI systems
Improvements to AI systems based on this paper:
-
Domain-aware architectural priors for small-sample, high-dimensional data: Replace generic CNN/Transformer backbones with a dual-pathway design that separates spectral (channel-grouped) and spatial (depthwise) feature extraction. This prevents premature entanglement of correlated features and reduces overfitting when training data is scarce (e.g., 560 samples). The improved system can achieve >97% accuracy with 5–18× fewer parameters than ResNet/ViT, making it deployable on edge devices.
-
Lightweight dual attention for selective emphasis without data hunger: Integrate a channel-wise squeeze-and-excitation (reduction ratio 4) coupled with a 1×1 convolution spatial gate, costing <5% of block parameters. Unlike self-attention, this stabilizes training under limited data and improves accuracy by 1–2% over no-attention baselines. The improved system can focus on informative spectral bands and spatial regions while remaining robust to noise and sample scarcity.
-
Hierarchical multi-scale encoding with controlled downsampling: Stack blocks across three stages with channel widths 64,128,256 and spatial resolutions 562,282,142, using layer-normalised strided convolutions for downsampling. This captures fine-to-coarse degradation cues (e.g., local tissue changes to global spoilage patterns). The improved system can classify freshness across 16 days with 0.64-day mean absolute error, with near-perfect accuracy (≥96% F1) in the commercially critical pre-expiry window (days 1–14).
-
Ordinal-aware evaluation and error concentration analysis: Use MAE and quadratic weighted kappa alongside accuracy to penalize large-distance errors. The improved system can identify that residual errors cluster in late spoilage (days 15–16) due to overlapping biological markers, enabling targeted retraining or rejection sampling for high-stakes decisions.
-
Leakage-controlled dataset curation for reproducible benchmarking: Adopt strict pack-level separation (train/val/test from distinct fish packs) to prevent data leakage. The improved system can be trained and evaluated on such datasets to yield trustworthy performance metrics, avoiding inflated accuracy from sample overlap.
-
Scalable throughput for real-time inspection: With 4.75M parameters and high throughput at batch sizes typical of conveyor belts, the improved system can run real-time hyperspectral classification on resource-constrained hardware (e.g., industrial PCs), with single-cube latency suitable for interactive point-of-acquisition feedback.
-
Ablation-guided hyperparameter selection: Use systematic ablations (e.g., grouped convolution groups g=8, kernel size 7, expansion ratios 2/4) to find an optimal accuracy–efficiency tradeoff. The improved system can be configured for different hardware budgets—e.g., shrinking channels to 48,96,192 for lower memory at modest accuracy loss, or deepening to 4,4,6 for better MAE at higher parameter cost.
-
Extensible framework for other food matrices: The architecture’s spectral-spatial factorization is not species-specific, so the improved system can be retrained on other hyperspectral food datasets (e.g., beef, poultry, produce) with minimal modification, provided similar spectral dominance and ordinal freshness labels exist.
What the improved AI system can do specifically:
-
Classify fish freshness from hyperspectral cubes with 97.8% accuracy and 0.64-day MAE, using only 4.75M parameters, in real-time on edge hardware.
-
Distinguish storage days with high reliability (≥96% F1) up to 8 days past the labelled expiry date, flagging only the most advanced spoilage stages (days 15–16) as uncertain.
-
Operate effectively with as few as 560 training samples, avoiding the data appetite of transformers.
-
Provide confidence-weighted predictions where late-stage errors are isolated, enabling automated quality control to reject only truly degraded products.
-
Be deployed in industrial inspection lines with 5–18× lower memory footprint than ResNet-50/ViT-B, reducing hardware costs while maintaining throughput.
Abstract
Hyperspectral imaging (HSI) offers nondestructive assessment of fish freshness by detecting biochemical alterations across spectral bands. However, conventional deep learning approaches do not fully address the particular characteristics of HSI data, such as spectral dominance over spatial textures, ordinal label structure, and a small number of training samples. We propose SGNet (Spectral-Grouped Network), a lightweight architecture that separates spectral and spatial feature extraction using grouped convolutions and a depthwise spatial pathway. A dual attention mechanism that couples channel-wise squeeze-and-excitation with spatial gating adaptively highlights informative features. SGNet achieves 97.8% classification accuracy and 0.64 days mean absolute error (MAE) with just 4.75M parameters when tested on our newly developed 16-day refrigerator-stored salmon fillet dataset. Ablation studies validate the contribution of each component, while comparisons demonstrate a five- to eighteen-fold parameter reduction relative to ResNet-50 and Vision Transformers. Our findings indicate that domain-aware design supports precise, real-time freshness prediction for industrial implementation.
Sources
Related papers
- Revisiting Integration of Image and Metadata for DICOM Series Classification: Cross-Attention and Dictionary Learning
- VesselSDF: Distance Field Priors for Vascular Network Reconstruction
- cSVR: Convolutional Slice-to-Volume Reconstruction
- NAIMA: Semantics Aware RGB Guided Depth Super-Resolution
- AneumoBench: A Source-Linked Benchmark for Synthetic-Geometry Transfer in Aneurysm CFD
- RETO: A Rotary-Enhanced Transformer Operator for High-Fidelity Prediction of Automotive Aerodynamics