MCLC-NET: Multimodal Continual Learning for Leaf Counting
cs.CV, cs.LG
Submitted: 2026-09-16
Updated: 2026-09-16
License: http://creativecommons.org/publicdomain/zero/1.0/
The gist: Leaf counting is an important task in plant phenotyping for monitoring plant growth and estimating crop yield.
Terminology
Abstract
Leaf counting is an important task in plant phenotyping for monitoring plant growth and estimating crop yield. Most existing methods rely on RGB images, but their performance is often affected by occlusion, lighting variations, and other real-world challenges. Additional modalities, such as depth and thermal images, can provide useful complementary information. However, multimodal leaf counting remains underexplored. Also, many existing methods assume that all training data are available simultaneously, which is impractical in real agricultural settings, where data is collected over time from multiple sources. To address these challenges, we propose MCLC-NET, a multimodal continual learning framework for leaf counting. It learns tasks sequentially using a memory-based strategy with a memory buffer to retain important samples from previous tasks. We also introduce MMLC, a real-world multimodal leaf-counting dataset designed for a domain incremental scenario (DIS) in CL. It contains RGB, depth, and thermal images collected across different crop types under varying environmental conditions, arranged in three orderings: crop-wise, time-wise, and mixed. Experimental results, averaged over three random seeds, demonstrate that MCLC-NET consistently outperforms existing methods across all three task orderings, achieving the lowest AMSE of 0.675 plus or minus 0.027, 0.542 plus or minus 0.069, and 0.745 plus or minus 0.057, respectively.
Sources
- Lifelong Learning with Dynamically Expandable Networks
- Online Continual Learning: A Systematic Literature Review of Approaches, Challenges, and Benchmarks
- Improving Object Counting with Heatmap Regulation
- Efficient Lifelong Learning with A-GEM
- Variational Prototype Replays for Continual Learning
- Dynamic Transformer Architecture for Continual Learning of Multimodal Tasks
- Don't forget, there is more than forgetting: new metrics for Continual Learning
- MUTEX: Learning Unified Policies from Multimodal Task Specifications
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models