Multi-perspective monitoring of wildlife and human activities from camera traps and drones with deep learning models

arXiv:2508.15629 · cs.CV · Submitted 2025-08-21 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Multi-perspective monitoring of wildlife and human activities from camera traps and drones with deep learning models".

Jane: Multi-perspective monitoring of wildlife and human activities from camera traps and drones with deep learning models addresses the need for understanding spatial distributions of wildlife and human activities to evaluate…

Tom: First, who's behind it and why it matters.

Title and authors: Tom: Wow, we're diving into this paper today, "Multi-perspective monitoring of wildlife and human activities from camera traps and drones with deep learning models." It sounds like they’ve put together a really cool system that combines two different ways of looking at the landscape.

Jane: That's right, Tom; it suggests using both ground-level camera traps for detailed observation and aerial drone imagery to get a broader view. It’s about making sure we see both what the animals are doing and what the people are up to in one comprehensive way.

Lu: The authors are from a solid group at UT Dallas, Auburn University, and San Diego State University; that tells you immediately this is coming from a place with strong expertise in geospatial science and ecology.

Meng: It’s interesting how they’ve brought together two very different data sources—fixed cameras versus moving drones—using deep learning to make sense of it all. I'm curious how much computational power this setup actually demands for real-time processing across different perspectives.

Lalam: From an AI standpoint, this approach is fascinating because it tackles the problem of spatial distribution by using a multi-modal input strategy, which is key for robust understanding in complex environments.

The paper's summary: Tom: So, what's the core idea behind this study? It’s about using these cameras and drones to figure out where wildlife and human activities are happening across a landscape.

Jane: Exactly, Tom; they are essentially building a system to map out human-wildlife interactions by using deep learning models to automatically identify those animals and people in the footage.

Lu: The paper summarizes how they integrated camera trap data with drone thermal imagery, using specific models like YOLOv11s for the cameras and an enhanced Faster R-CNN for the drones, which is a pretty sophisticated setup.

Meng: That's where I get practical; it sounds like they are trying to overcome the limitations of relying on just one sensor type by having this multi-perspective view. Does this mean they can see things that one sensor would miss?

Lalam: The paper highlights how integrating these different data streams allows for a richer understanding of the conflict zones because you get both ground-level detail and aerial context simultaneously.

The paper's improvements: Tom: Moving into what they actually achieved, this study focuses on using YOLOv11s for camera traps and an enhanced Faster R-CNN model for drone thermal imagery to detect targets.

Jane: They used these specific models because the results showed that YOLOv11s performed really well in detecting objects in camera trap images, achieving a precision of ninety-six point two percent and a recall of ninety-two point three percent.

Lu: And for the drone side, they employed an enhanced Faster R-CNN model with FPN and ResNet18 to detect deer, getting an average precision score of ninety-one point six percent across all deer objects in their thermal imagery analysis.

Meng: From an engineering view, seeing those specific performance metrics is helpful because it gives us a benchmark for what the system can reliably do when dealing with challenging visual data from both sources.

Lalam: This level of detail shows how tailored deep learning models can significantly boost detection accuracy compared to using a single type of sensor or model alone.

Conclusion: Tom: So, wrapping up this paper, the main point is that combining these multi-perspective monitoring techniques really helps reveal human–wildlife conflicts in conserved landscapes.

Jane: It shows that using automated object detection from both camera traps and drones creates a much more effective way to monitor wildlife and manage those interactions spatially.

Lu: The integration of this multi-perspective monitoring framework offers a promising structure for developing future three-dimensional and flexible wildlife monitoring systems in protected areas.

Meng: I think the real impact here is providing actionable data; if you can automatically map conflict zones, conservation managers can start deploying resources where they are most needed.

Lalam: This work really demonstrates how sophisticated AI methods can be used to provide the kind of detailed spatial intelligence that local authorities need to develop targeted management strategies for wildlife and human interactions.

Geospatial Information Science, the University of Texas at Dallas

cs.CV

Submitted: 2025-08-21

Updated: 2025-08-21

Journal ref: Remote Sens. 2026, 18, 3375

DOI: 10.3390/rs18193375

License: http://creativecommons.org/licenses/by-nc-sa/4.0/

Importance score: 82/100

The gist: Multi-perspective monitoring of wildlife and human activities from camera traps and drones with deep learning models addresses the need for understanding spatial distributions of wildlife and human

Key concepts

Multi-perspective Monitoring
This involves using different sources—like camera traps and drones—to observe the same area from multiple angles simultaneously. This comprehensive view is essential for accurately understanding where wildlife and humans are interacting, rather than relying on just one type of observation.
YOLOv11s
YOLOv11s is a specific deep learning model used to automatically detect objects in camera trap images. It was highly effective, achieving high precision and recall, making it the best tool for quickly identifying animals and humans within the collected camera trap photos.
Kernel Density Estimation (KDE)
KDE is a spatial analysis technique used to map out where certain events or objects are most concentrated in an area. By applying this to wildlife detections and human activity detections, researchers could pinpoint specific hotspots where wildlife and people frequently overlap, indicating conflict areas.

Terminology

Summary

Multi-perspective monitoring of wildlife and human activities from camera traps and drones with deep learning models addresses the need for understanding spatial distributions of wildlife and human activities to evaluate interactions and inform conservation planning. The gist: This study reveals human–wildlife conflicts within the conserved landscape by integrating multi-perspective monitoring with automated object detection, using YOLOv11s for camera trap imagery and an enhanced Faster R-CNN model for drone thermal imagery.

Context

Wildlife and human activities are key components of landscape systems, and understanding their spatial distribution is essential for evaluating human–wildlife interactions and informing effective conservation planning. Human activities frequently change or destroy wildlife habitats, with illegal hunting being a detrimental impact. Wildlife interacts with humans in varying frequencies, leading to positive or negative interactions referred to as human-wildlife conflict. Monitoring these activities is crucial for understanding these interactions and playing a role in effective wildlife conservation and sustainable ecological landscape management.

Methods

The study was conducted in Chitwan National Park (CNP), Nepal, and adjacent regions between February and July 2022. Images were collected by visible and near-infrared camera traps, processed to create training and testing datasets for deep learning models to automatic identify wildlife and human activities. Drone-collected thermal imagery was used for detecting targets to provide a multiple monitoring perspective. The dataset included images of domestic goats, tigers, rhinos, rhesus macaques, elephants, spotted deer, sambar deer, and humans. To create the dataset from camera trapping images containing both wildlife and human activities (Table 1), the Roboflow platform was utilized to annotate images accurately for YOLO model efficiency. The drone imagery was captured using a DJI Mavic 2 Enterprise equipped with both thermal and optical visual cameras, capturing infrared and visible RGB images simultaneously.

Deep Learning Models

The study tested several deep learning models for object detection in camera trap imagery. YOLOv11s achieved the highest performance with a precision of 96.2%, recall of 92.3%, mAP50 of 96.7%, and mAP50-90 of 81.3, making it the most effective for detecting objects in camera trap imagery. YOLO models were utilized, including YOLOv5n, YOLOv5s, and YOLOv5m, along with YOLOv11n, YOLOv11s, and YOLOv11m. The training process utilized an input image size of 640 and 32 as the batch size for the models. For drone-based thermal imagery analysis, an enhanced Faster R-CNN model was used to detect deer, achieving an Average precision score of 91.6% for all deer objects when integrated with FPN and ResNet18.

Spatial Analysis

Spatial pattern analysis was performed to identify animal and resident activity hotspots and delineation potential human–wildlife conflict zones. Kernel Density Estimation (KDE) was employed to identify hotspots where wildlife and human activities intersect, using results from camera trapping detection for wildlife (tigers, rhinos) and detections of people and domestic goats for human activities. The Kernel Density spatial analysis tool in ArcGIS Pro was used to calculate density by applying a Kernel function to create a continuous surface. Clustering results based on correlation-based hierarchical clustering categorized the 80 camera traps into three distinct groups: human, wildlife, and conflict zones, where both humans and wildlife are frequently detected by the same camera traps.

Conclusions

This study reveals human–wildlife conflicts within the conserved landscape. Integrating multi-perspective monitoring with automated object detection enhances wildlife surveillance and landscape management. The integration of camera traps and drone-based monitoring provides a promising framework for future three-dimensional and flexible wildlife monitoring systems, offering actionable data to inform local authorities in the development of targeted management strategies. Future research will aim to advance this framework by incorporating automated individual identification and analyzing spatial interactions using state-of-the-art machine learning and geographic information system (GIS) techniques within protected landscapes.


The gist

This study reveals human–wildlife conflicts within the conserved landscape by integrating multi-perspective monitoring with automated object detection, using YOLOv11s for camera trap imagery and an enhanced Faster R-CNN model for drone thermal imagery.

How it works

  1. Camera trapping data were collected throughout the study area to support the training and detection of deep learning models, capturing images with motion-sensitive camera traps that produce motion-activated 32MP photos and 4K 30 fps videos. These systems are strategically positioned at sites exhibiting signs of wildlife and human activities, with GPS coordinates recorded for each trap location.

  2. To reflect human activities, images including human and domestic goats were collected in the dataset, alongside images of wildlife such as spotted deer, sambar deer, rhesus macaques, rhinos, tigers, and elephants.

Improvements for AI systems

Here are specific improvements to AI systems based on the provided research, along with what those improved systems could achieve:


  1. [Improved System]: A Multi-Modal Surveillance Platform integrating Camera Trap (RGB/NIR) and Drone Thermal Imagery via a Federated Deep Learning Architecture.

  2. [Capability]: This system can perform real-time, cross-perspective detection of wildlife and human activities across large, complex landscapes (e.g., national parks). Specifically:

  3. [Specific Improvement 1 - Detection Accuracy]: By integrating YOLOv11s (for camera trap imagery) with an enhanced Faster R-CNN model (for thermal drone imagery), the system achieves superior target detection accuracy compared to single-sensor systems, particularly under challenging conditions like low light, heavy vegetation occlusion, and partial body visibility.

  4. [Specific Improvement 2 - Data Fusion]: The system can perform effective data fusion by cross-validating detections between ground-level camera traps and aerial thermal imagery (e.g., confirming a detected animal presence from a drone flight path against ground evidence), thereby reducing false positives and increasing the reliability of abundance estimates.

  5. [Specific Improvement 3 - Automated Conflict Zoning]: The system can automatically generate dynamic spatial maps identifying high-probability human-wildlife conflict zones by performing Kernel Density Estimation (KDE) on aggregated detection results from both data sources, specifically targeting overlaps between human activity classes (people, domestic goats) and key wildlife species (tigers, rhinos).

  6. [Specific Improvement 4 - Predictive Modeling]: The system can utilize the spatial clustering results derived from hierarchical clustering to categorize camera trap locations into distinct ecological groups (Human-dominated, Wildlife-dominated, Conflict zones), allowing conservation managers to proactively predict where future conflicts are most likely to occur based on current spatial patterns.

  7. [Specific Improvement 5 - Species/Activity Identification]: The system can be trained to automatically identify specific wildlife species and human activities from raw imagery using state-of-the-art deep learning models (like YOLOv11), moving beyond simple object detection to detailed ecological classification for precise monitoring.

  8. [Improved System]: A Real-Time, Adaptive Wildlife Monitoring Network utilizing Multi-Scale Object Detection Models on Unmanned Aerial Vehicles (UAVs).

  9. [Capability]: This system can enable flexible, adaptive aerial surveillance by leveraging the YOLO architecture (specifically YOLOv5/YOLOv11 variants) optimized for real-time performance on drone imagery. Specifically:

  10. [Specific Improvement 6 - Real-Time Aerial Tracking]: The system can perform high-speed, real-time detection and tracking of specific wildlife species (e.g., deer, rhinos) using thermal imagery from drones, allowing for rapid response to unusual movements or incursions into sensitive areas.

  11. [Specific Improvement 7 - Enhanced Thermal Differentiation]: By utilizing thermal infrared sensors on drones, the system can effectively detect targets obscured by vegetation or operating under low-light conditions where standard optical systems fail, providing a critical 'eyes in the dark' capability for monitoring nocturnal wildlife or hidden human movements.

Abstract

Wildlife and human activities are key components of landscape systems. Understanding their spatial distribution is essential for evaluating human wildlife interactions and informing effective conservation planning. Multiperspective monitoring of wildlife and human activities by combining camera traps and drone imagery. Capturing the spatial patterns of their distributions, which allows the identification of the overlap of their activity zones and the assessment of the degree of human wildlife conflict. The study was conducted in Chitwan National Park (CNP), Nepal, and adjacent regions. Images collected by visible and nearinfrared camera traps and thermal infrared drones from February to July 2022 were processed to create training and testing datasets, which were used to build deep learning models to automatic identify wildlife and human activities. Drone collected thermal imagery was used for detecting targets to provide a multiple monitoring perspective. Spatial pattern analysis was performed to identify animal and resident activity hotspots and delineation potential human wildlife conflict zones. Among the deep learning models tested, YOLOv11s achieved the highest performance with a precision of 96.2%, recall of 92.3%, mAP50 of 96.7%, and mAP50 of 81.3%, making it the most effective for detecting objects in camera trap imagery. Drone based thermal imagery, analyzed with an enhanced Faster RCNN model, added a complementary aerial viewpoint for camera trap detections. Spatial pattern analysis identified clear hotspots for both wildlife and human activities and their overlapping patterns within certain areas in the CNP and buffer zones indicating potential conflict. This study reveals human wildlife conflicts within the conserved landscape. Integrating multiperspective monitoring with automated object detection enhances wildlife surveillance and landscape management.

Sources

Related papers