DeepForestVisionV2: Ecology-Driven Taxonomy Expansion for Camera-Trap Monitoring in African Tropical Forests
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "DeepForestVisionV2: Ecology-Driven Taxonomy Expansion for Camera-Trap Monitoring in African Tropical Forests".
Tom: Camera-trap monitoring in African tropical forests is being enhanced by DeepForestVisionV2, an ecology-driven expansion that addresses deployment challenges across vertical stratification, scene openness, and anthropogenic interfaces.
Jane: First, who's behind it and why it matters.
Title and authors: Tom: Okay, we're starting with the title and authors of "DeepForestVisionV2: Ecology-Driven Taxonomy Expansion for Camera-Trap Monitoring in African Tropical Forests." It’s a very descriptive name that tells us exactly what they’re tackling: using ecological needs to guide how the AI classifies wildlife in these specific forest settings.
Jane: I think the title immediately signals that this isn't just about making the existing system bigger; it’s about tailoring it to fit different situations where cameras are actually placed. It suggests a shift from a general model to one that understands local conditions.
Lu: The authors, with their affiliations spanning ecology and computer science centers, give us confidence that they have the right mix of expertise to tackle both the biological complexity and the deep learning architecture needed for this kind of expansion.
Meng: I'm just thinking about how much data they used to train this thing; if it’s been trained on a massive dataset from multi-country projects, that suggests a solid foundation, but I want to know how robust it is when things get really messy in the field.
Lalam: The authors’ work shows a deep commitment to bridging the gap between raw detection and meaningful ecological interpretation, which is something we need as AI moves further into applied science.
The paper's summary: Tom: Now for a quick rundown of what this paper actually says about DeepForestVisionV2. Essentially, they take the original model, which worked well for closed-canopy ground-level setups, and they significantly increase the prediction classes from thirty-five to sixty-four. This expansion is directly tied to addressing three specific deployment challenges: vertical stratification, scene openness, and human interfaces like park edges.
Jane: That’s a big jump in complexity, Tom. So, instead of just labeling everything as a general animal or bird, they are adding classes specifically for things you see when the camera is near clearings or at riverbanks. It’s about making the labels more specific to where the camera actually is.
Lu: The authors explain that this expansion includes sixty-one animal classes alongside human, vehicle, and blank categories, which directly targets those deployment gradients they mentioned—the vertical change in canopy visibility, the openness of the scene near water, and distinguishing between wildlife and domestic animals at park edges.
Meng: So it’s not just adding more labels for fun; it’s a direct response to field realities where simple labels like "monkey" aren't enough to tell a park manager what management action to take. That makes the practical utility much clearer.
Lalam: It really shows how refining the output space based on real ecological contexts can dramatically improve the actionable information we get from vision systems in conservation efforts.
The paper's improvements: Tom: Let’s talk about these specific improvements they’ve implemented, which is where the real technical substance of DeepForestVisionV2 lies. They kept the original architecture for generating one label per camera-trap item, but they fine-tuned the classification stage using a DINOv3 ViT-B/sixteen model on crops from MegaDetector v5 to create those sixty-four classes.
Jane: The methodology involves applying MegaDetector v5 to find instances, and then sending the animal detections into the DINOv3 classifier for more granular labeling. They also used specific augmentation techniques like color jitter and Gaussian blur during training, which is smart for making the model robust against varying field conditions.
Lu: The results show that this expansion is not just theoretical; they tested it across different scenarios, and in the forest interior benchmark, DeepForestVisionV2 increased the number of surfaced taxa from twenty-two to twenty-nine. They also saw similar increases at other sites, like riverbanks where it went from four to nine identified taxa.
Meng: That improvement from twenty-two to twenty-nine is a solid metric because it shows that the model isn't just guessing; it’s actually finding more relevant species when the context changes, which is what we need for reliable deployment.
Lalam: The paper also pointed out that at the park-edge benchmark, this system improved accuracy from zero point six two to a very solid zero point eight six and, critically, it managed to remove all eleven false alarms while keeping the recall high.
Conclusion: Tom: So we’re wrapping up with the conclusions of "DeepForestVisionV2: Ecology-Driven Taxonomy Expansion for Camera-Trap Monitoring in African Tropical Forests." The main point is that this expanded system provides material improvements to field utility by aligning the classification space with the three distinct ecological gradients—vertical, openness, and human interfaces.
Jane: It really boils down to making sure the AI output matches what ecologists actually need when they are out in the field, giving them much finer detail on primates and riparian communities.
Lu: The implication for future work seems to be building upon this foundation by using these finer-grained labels for more complex behavioral analyses or even predictive modeling about community shifts along those gradients.
Meng: From an engineering standpoint, the robustness of maintaining an offline workflow for both photographs and videos means this tool has a good chance of being adopted by field teams who can’t rely on constant cloud access. That sustained offline capability is a key practical factor here.
Lalam: For culture, I see this as showing that deep domain knowledge—in this case, ecological knowledge—can directly inform the design of powerful AI tools to solve tangible conservation problems effectively.
Tom: It’s a lot of exciting stuff showing how targeted model expansion can yield very specific and useful results for monitoring wildlife in challenging tropical environments. We’ve covered a lot about DeepForestVisionV2 today.
Hugo Magaldi, Théau d’Audiffret, Etienne François Akomo-Okoue, Bala Amarasekaran, Naomi Anderson, Claire Auger, Noémie Cappelle, Daniel Cornélis, Raphaël Cornette, Tobias Deschner, Gabriel Dubus, Davy Fonteyn, Rosa M. Garriga, Jennifer Hatlauf, Innocent Kasekendi, Raymond Katumba, Aram Kazandjian, Alfred Ngomanda, Stéphan Ntie123 Simone Pika10 Xavier Rufray, Harold Rugonge, John Justice Tibesigwa14 Peter van Lunteren15 Hadrien Vanthomme7 Joeri A. Zwerts16 and Sabrina Krief
UMR7206 Eco-Anthropologie, MNHN, Paris, France · One Forest Vision initiative Sebitoli Chimpanzee Project, Sebitoli, Kibale National Park Uganda · Centre National de la Recherche Scientifique et Technologique Libreville, Gabon · Institut de Recherche en Ecologie Tropicale Libreville, Gabon · Tacugama Chimpanzee Sanctuary Freetown Sierra Leone · International Department Biotope Montpellier, France · CIRAD UPR Forêts et Sociétés Montpellier, France · Institut de Systématique Evolution Biodiversité ISYEB UMR7205 CNRS MNHN SU EPHE-PSL UA Paris, France · Max Planck Institute for Evolutionary Anthropology Leipzig, Germany · Comparative BioCognition Institute of Cognitive Science Osnabrück University Osnabrück, Germany · BOKU University Institute of Wildlife Biology and Game Management Vienna, Austria · Uganda Wildlife Authority Kampala, Uganda · Addax Data Science Utrecht, The Netherlands · Wildlife Ecology and Nature Restoration Utrecht University, The Netherlands
cs.CV, q-bio.QM
Submitted: 2026-06-18
Updated: 2026-09-29
Comments: Published in Pattern Recognition. ICPR 2026 International Workshops (LNCS 17113, pp. 252-265). Please cite the published version: https://doi.org/10.1007/978-3-032-39518-4_17
Journal ref: In: Sintorn, IM., Pal, U., Lu, S. (eds) Pattern Recognition. ICPR 2026 International Workshops. ICPR 2026. Lecture Notes in Computer Science, vol 17113. Springer, Cham
DOI: 10.1007/978-3-032-39518-4_17
Code: https://github.com/MNHN-OFVI/DeepForestVisionV2
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 83/100
The gist: Camera-trap monitoring in African tropical forests is being enhanced by DeepForestVisionV2, an ecology-driven expansion that addresses deployment challenges across vertical stratification, scene
Key concepts
- Vertical Stratification
- This refers to different layers within a forest environment where animals are found at varying heights. DeepForestVisionV2 was expanded to better identify primates visible near clearings or canopy openings, which were previously missed by models focused only on ground-level deployments.
- Openness Gradient
- This gradient describes environments that are more exposed, such as riverbanks. The expanded model helps capture detections of birds and semi-aquatic animals in these open views, improving monitoring accuracy in areas with less dense forest cover.
- Anthropogenic Interface
- This relates to areas where wildlife interacts with human activities, like park edges. The new classes allow the system to accurately distinguish between wildlife and domestic animals, which is crucial for management actions at these boundaries.
Terminology
Summary
Camera-trap monitoring in African tropical forests is being enhanced by DeepForestVisionV2, an ecology-driven expansion that addresses deployment challenges across vertical stratification, scene openness, and anthropogenic interfaces. The core finding is that expanding the prediction space from 35 to 64 classes significantly improves field utility by better capturing taxa relevant to diverse ecological gradients while maintaining a robust offline workflow for both photographs and videos.
The Gist
DeepForestVisionV2 expands DeepForestVision from a 35-class forest-interior camera-trap model to a 64-class system better aligned with field realities in African tropical forests.
Motivation and Problem Statement
Camera traps are crucial for wildlife monitoring, but their ecological value depends on deployability under field constraints and taxonomic resolution. Previous tools like DeepForestVision were designed around closed-canopy, ground-level deployments. However, ongoing projects increasingly deploy cameras along ecological gradients: (i) a vertical gradient where arboreal primates become visible near clearings or canopy openings; (ii) an openness gradient where riverbanks expose birds and semi-aquatic taxa; and (iii) an anthropogenic gradient where park-edge cameras must distinguish wildlife from livestock. In these regimes, the original coarse labels are often insufficient for ecological interpretation or management action.
DeepForestVisionV2 addresses this mismatch by redesigning the label space to better capture taxa relevant along these deployment gradients.
Taxonomy Expansion and Ecological Motivation
DeepForestVisionV2 expands the prediction space from 35 to 64 classes, which includes 61 animal classes plus human, vehicle, and blank.
This expansion is motivated by three recurrent deployment gradients:
-
Vertical stratification: To capture taxa like
arboreal primates become visible near clearings, mineral licks, and canopy openings.
-
Openness gradient: To expose detections of
birds [at] riverbanks and open views
and semi-aquatic taxa. -
Anthropogenic interface: To allow cameras to distinguish wildlife from domestic animals at park edges.
The new taxonomy provides finer-grained downstream ecological analyses, increasing DeepForestVision’s native 35-class resolution for arboreal primates, birds, riparian taxa, and domestic species (Table 3). For instance, the expansion includes new classes such as chimpanzee,
gorilla,
hippopotamus,
and specific primate refinements like red-capped mangabey.
Methodology: Pipeline and Training
DeepForestVisionV2 retains the original DeepForestVision architecture to produce one label per camera-trap item (photograph or video). The pipeline involves two stages: first, MegaDetector v5 is applied to detect animal, human, and vehicle instances. Human and vehicle detections are assigned directly from detector outputs, while animal detections are cropped and sent to a DINOv3 ViT-B/16 classifier. Class scores are computed on retained detections across sampled frames and then averaged to produce one item-level prediction. If no detection remains after thresholding, the item is labeled blank.
The classification stage was fine-tuned on MegaDetector crops using DINOv3 ViT-B/16 with specific augmentation techniques, including random resized crop, color jitter, random grayscale, and Gaussian blur. The training ran for 20 epochs with checkpoint selection based on validation balanced accuracy.
Evaluation and Results
Evaluation is split into two parts: a cross-country cropped-photo validation set to assess robustness across sites and settings (where DeepForestVisionV2 reached 0.86 accuracy), and three held-out Uganda video benchmarks corresponding to the three targeted ecological gradients.
Performance on deployment benchmarks shows improvements:
= Forest interior benchmark:
DeepForestVisionV2 matches the original DeepForestVision in accuracy (0.89 for both models) while increasing the number of surfaced taxa from 22 to 29.
= Riverbanks benchmark:
DeepForestVisionV2 remains close in accuracy (0.72 vs. 0.73) and raises the number of surfaced taxa from 4 to 9.
= Park-edge benchmark:
DeepForestVisionV2 improves accuracy from 0.62 to 0.86, and critically, removes all false alarms while keeping high recall.
Ecological Utility and Calibration
To quantify ecological gain, the paper reports the number and identity of taxa surfaced by DeepForestVisionV2 but not represented in DeepForestVision. For the park-edge benchmark involving primates and goats, a simple alarm task showed that DeepForestVisionV2 removes all 11 false alarms.
Regarding calibration, pooled video-benchmark predictions show an Expected Calibration Error (ECE) of 0.17.
Improvements for AI systems
Here are specific improvements for AI systems based on the DeepForestVisionV2 methodology, and what those improved systems can achieve:
-
The core improvement is transitioning from a fixed, coarse taxonomy (35 classes) to an ecology-driven, context-aware classification space of 64 classes.
-
The system can now perform high-resolution ecological monitoring across three critical deployment gradients:
-
When deployed in forest interiors, the system can accurately identify and count specific arboreal primates (e.g., red colobus, grey-cheeked mangabey) and specialized bird groups (e.g., crane, francolin), allowing for detailed population dynamics studies that were previously obscured by coarse labels like
monkey
orbird.
-
When deployed near riverbanks or open views, the system can reliably detect semi-aquatic taxa (e.g., otter) and specific bird guilds (e.g., crane, duck), enabling accurate assessment of riparian community composition in open landscapes.
-
When deployed at park edges or interfaces, the system can specifically distinguish between wildlife and anthropogenic confounders like domestic goats and cattle with high accuracy (0.86 validation accuracy reported), which directly translates into operational utility by suppressing false alarms from livestock passages (reducing false alarms from 11 to 0 in the park-edge use case).
-
The improved system maintains a robust, offline workflow for both photographs and videos, ensuring deployability in resource-constrained field settings without requiring constant cloud connectivity or complex online processing.
-
The system provides quantitative metrics for ecological gain: researchers can directly quantify the number and identity of taxa surfaced by the expanded taxonomy but not present in the original model, allowing for direct comparison between baseline and enhanced monitoring effectiveness.
-
For automated alerting pipelines, confidence scores from DeepForestVisionV2 can be interpreted with caution (as noted by its Expected Calibration Error of 0.17), suggesting that post-hoc calibration techniques (like temperature scaling) must be applied before using confidence thresholds for critical triage or alarming, leading to more reliable operational decision-making.
Abstract
Camera-trap monitoring in African tropical forests increasingly extends beyond closed-canopy interiors to riverbanks, clearings, and park edges. Among available open tools for African forest camera-trap classification, DeepForestVision is the only one providing a matched offline workflow for both photographs and videos, and previous work showed that it outperformed other available baselines on a comparable benchmark. However, it was designed for closed-canopy, ground-level forest interiors and uses a 35-class prediction space that becomes too coarse when deployments encounter arboreal primates, birds, semi-aquatic taxa, or human-associated confounders such as livestock. We present DeepForestVisionV2, an ecology-driven expansion from 35 to 64 prediction classes (61 animal classes plus human, vehicle, and blank) designed to address three recurrent deployment gradients: vertical stratification, scene openness, and anthropogenic interfaces. DeepForestVisionV2 retains the same offline workflow and is trained on 1,535,010 photographs and 243,354 videos from multi-country African tropical-forest projects. Evaluation combines a cross-country cropped-photo validation set, used to assess robustness across sites and camera-trap settings, with three held-out Uganda video benchmarks spanning the targeted gradients. On the validation set, DeepForestVisionV2 reaches 0.86 accuracy, 0.82 macro-F1, and 0.81 balanced accuracy. On the deployment benchmarks, it preserves or improves baseline accuracy despite its harder classification task, while increasing the number of identified taxa from 22 to 29 in forest-interior videos and from 4 to 9 at riverbanks. In the park-edge use case, it raises accuracy from 0.62 to 0.86 and reduces false alarms from 11 to 0. These results show that DeepForestVisionV2 materially improves field utility while preserving robustness across sites, habitats, and camera-trap settings.
Sources
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models