Mapping and Classification of Trees Outside Forests using Deep Learning

arXiv:2510.25239 · cs.CV · Submitted 2025-10-29 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Mapping and Classification of Trees Outside Forests using Deep Learning".

Jane: Trees Outside Forests (TOF) play an important role in agricultural landscapes by supporting biodiversity, sequestering carbon, and regulating microclimates.

Tom: First, who's behind it and why it matters.

Title and authors: Tom: So, we're getting into a paper called "Mapping and Classification of Trees Outside Forests using Deep Learning." It sounds pretty technical at first, but basically, the researchers are looking at how we can use deep learning to figure out where trees exist outside of established forests in agricultural areas.

Jane: That’s right, Tom. The authors are Moritz Lucasa and his team from Osnabrück University and the Leibniz Institute for Agricultural Engineering and Bioeconomy. Their whole goal is to move away from those old, rigid rule-based methods that often struggle when you look at different types of farmland across Germany.

Lu: It’s really interesting because they’re directly tackling a problem where traditional methods, like those relying on simple threshold rules for height or spectral data, just don't adapt well to the varied landscapes we see everywhere.

Meng: From an engineering standpoint, I wonder if this deep learning approach can actually handle the complexity of real-world agricultural imagery without needing impossibly clean training data.

Lalam: I think what’s exciting here is that they are moving from simple binary detection to a four-class classification system: Forest, Patch, Linear, and Tree.

Tom: Exactly! They are trying to give us a much finer detail of the landscape than just saying "there's forest or there isn't," which is what they call the TOF classification problem.

Jane: And they’re testing this approach across four very different agricultural settings in Germany—from hedgerow structures to intensive farming plots—which sets up a really strong test for how well the model actually works in practice.

Lu: Testing spatial generalization across those distinct regions is crucial because it shows whether a model learned just one specific pattern or if it truly understands the underlying ecological structures, which is a big step forward for agricultural mapping.

Meng: I’m curious how they handled the data collection across those different environments, since that’s usually where models break down in deployment scenarios.

The paper's summary: Tom: Okay, so what they actually did is present their new dataset and then compare six different types of deep learning architectures—CNNs, vision transformers, and hybrid models—to see which one does the best job mapping those trees outside forests.

Jane: They built this new dataset using high-resolution aerial imagery and digitally derived surface models from four German landscapes, focusing on capturing a wide range of landscape structures in one place.

Lu: The core contribution is comparing these different architectures—like FT-UNetFormer versus U-Net—to pinpoint which one is most effective for this specific task, which moves beyond just using a single model type.

Meng: It sounds like they aren't just proving a model works; they’re systematically vetting the entire class of architectures against each other to find the optimal structure for this type of spatial segmentation.

Lalam: From my perspective, it’s important that they are looking at hybrid models, because combining local feature extraction with global context modeling seems like a smart way to handle both fine details and larger structures simultaneously.

Tom: And the summary points out that FT-UNetFormer came out on top, achieving mean metrics of zero point seven three nine for the mean IoU score and zero point eight four three for the mF1 score.

Jane: That result is solid, showing that this new deep learning method actually outperforms the previous approaches they were comparing against in terms of overall accuracy across all the classes.

Lu: The paper emphasizes that their analysis of accuracy across TOF classes and landscape structures helps us see exactly where a model is strong and where it has weaknesses, which is essential for real-world application.

Meng: So they’re not just giving us a high number; they’re giving us the map of where the model succeeds and fails when dealing with different kinds of land use, which is more useful for engineers than just a single score.

The paper's improvements: Tom: Now, let's talk about what the authors suggest as improvements for this work itself, because they aren't just stopping there with the results.

Jane: They suggest focusing on architectural advancements, specifically favoring vision transformers over purely CNN-based models because they seem better at capturing that long-range spatial context needed for these complex TOF patterns.

Lu: I agree with Jane; the paper shows that integrating transformer elements into a model, like in FT-UNetFormer, seems to give it an advantage in understanding how distant landscape features relate to each other.

Meng: From a practical standpoint, I’m interested in the idea of hybrid architectures; combining CNNs for local detail with transformers for global context sounds like a very robust way to build something that actually performs well on diverse imagery.

Lalam: I think they also suggest using Swin Transformer backbones, which are known for efficiently capturing multi-scale features, so incorporating that into the design could really enhance the model's ability to see different sizes of TOF elements.

Tom: And they also touch on data strategy improvements, like curating training sets from regions with similar structures when you plan to deploy this model in a new area.

Jane: That’s important because it shows that structural similarity between training data and the target environment drives how well the model can generalize, which is a key point for any deployment strategy.

Lu: I think if they could get better at mitigating class imbalance, especially for those smaller classes like 'Patch', that would be a big step toward making the classification more reliable in real-world scenarios.

Meng: And on the deployment side, the authors hint at using sliding window approaches with overlapping predictions to resolve boundary ambiguities, which I think is a necessary practical refinement for getting high-quality maps.

Conclusion: Tom: So, wrapping up this discussion on the "Mapping and Classification of Trees Outside Forests using Deep Learning," the main conclusion is that FT-UNetFormer achieved the best overall results with a mean IoU of zero point seven three nine and an mF1 score of zero point eight four three.

Jane: That’s a strong performance metric, Tom, especially since they tested it across those four distinct German agricultural landscapes, which shows the method is quite adaptable.

Lu: It really confirms that the combination of vision transformers and CNN structures provides a solid framework for tackling this kind of complex spatial problem when you have varied input data.

Meng: For practical impact, it means we can start thinking about automated tools for land-use planning that provide this level of fine detail without needing constant manual inventory work.

Lalam: It’s exciting because if we get models that can accurately map these features, it could actually help inform how we manage carbon sequestration and biodiversity in our agricultural regions.

Tom: Exactly! We've seen the results on this paper, "Mapping and Classification of Trees Outside Forests using Deep Learning," which shows a very capable system for handling these kinds of challenges.

Jane: It really gives us a solid foundation to see how these deep learning architectures can be applied to many other challenging environmental mapping tasks we face.

Lu: I think the future work should focus on making sure the model is robust enough for deployment in those highly specific, real-world farming contexts we discussed earlier.

Meng: From an engineering side, I’m looking forward to seeing how this framework scales up when we move from testing tiles to mapping entire regions with high fidelity.

Lalam: And if we can refine these models, they could significantly improve the way AI supports our cultural and environmental goals by providing detailed, actionable spatial data.

Osnabrück University · Leibniz Institute for Agricultural Engineering and Bioeconomy

cs.CV

Submitted: 2025-10-29

Updated: 2026-10-05

Code: https://github.com/Moerizzy/TOFMapper

Importance score: 86/100

The gist: Trees Outside Forests (TOF) play an important role in agricultural landscapes by supporting biodiversity, sequestering carbon, and regulating microclimates.

Key concepts

Trees Outside Forests (TOF)
These are non-forest tree areas within agricultural landscapes that are important for biodiversity and carbon storage. The study defined them using geometric rules based on area and shape ratios derived from Digital Orthophotos and Digital Surface Models.
Reference Data Generation
The ground truth for the classification was created by filtering Digital Surface Models to select vegetation areas, calculating the Normalized Difference Vegetation Index (NDVI), and using unsupervised clustering. This process refined geometric classes like Forest, Patch, Linear, and Tree.
FT-UNetFormer
This is a specific deep learning architecture tested in the study. It was found to be the most effective model for mapping TOF because it achieved the highest mean Intersection-over-Union (mIoU) score of 0.739 and mean F1 score of 0.843 across all testing tiles.

Terminology

Summary

Trees Outside Forests (TOF) play an important role in agricultural landscapes by supporting biodiversity, sequestering carbon, and regulating microclimates. The gist: FT-UNetFormer achieved the best mean metrics, with mIoU score of 0.739 and mF1 score of 0.843.

Study Areas and Data Acquisition

The study utilized four distinct agricultural landscapes in Germany to test the methodology's spatial generalization across different environments. These areas were selected to capture a range of landscape structures, with each site having a unique character:

  1. SchleswigHolstein (SH): Distinguished by low tree cover and a high proportion of TOF, characterized by hedgerow structure and traditional smallholder farming practices.

  2. Brandenburg (BB): Dominated by very large agricultural plots shaped by collectivized agriculture, resulting in the lowest TOF rate among the study areas.

  3. North Rhine-Westphalia North (NRW N): Represents a smallholder agricultural landscape with intensive farming.

  4. North Rhine-Westphalia South (NRW S): A more hilly region with elevation ranging from 300 to 500 m above sea-level, characterized by extensive grasslands and forests.

The database consisted of Digital Orthophotos (DOP) and photogrammetrically derived normalized Digital Surface Models (nDSM), captured on the same dates between 2021 and 2023. All DOP were collected during the summer months (June - August) and included four spectral bands: RGB and near-infrared. The raster data were resampled to a spatial resolution of 20 cm using nearest-neighbor interpolation.

Reference Data Generation

The TOF classification relies on four geometric classes derived from previous studies, which were refined through an automated approach with manual refinement. The classes are defined as:

(See Figure 2 for class definitions)

  1. Forest: Wooded areas that are larger than 5,000 m2 and elongation ratio (length/width) smaller than 3.

  2. Patch: Areas between 500 m2 and 0.5 ha with an elongation ratio smaller than 3.

  3. Linear: Areas with an elongation ratio greater than 3.

  4. Tree: Areas smaller than 500 m2.

The reference data was created by first generating a mask filtering nDSM values below a threshold of 3 m. Within this mask, the Normalized Difference Vegetation Index (NDVI) was calculated, and an unsupervised k-means clustering with k = 2 was applied to split pixels into active vegetation and the other impervious surfaces. Small gaps within the trees mask were filled using a morphological closing operation with a 5 × 5 pixel structuring element. The Douglas–Peucker algorithm was then applied to simplify and smooth the edges while preserving the overall shape and maintaining important details. Finally, polygons were classified based on FAO definitions: All polygons with a width bigger than 20 m and an area bigger than 0.5 ha were classified as Forests and the remaining into the three TOF classes.

Model Selection and Training Setup

The study systematically compared six state-of-the-art semantic segmentation architectures to identify the most effective model for TOF mapping. The models evaluated included: ABCNet, BANet, DC-Swin, FT-UNetFormer, LSKNet, and U-Net. The training was conducted on four NVIDIA A100 GPUs using a batch size of 8 (limited to 4 for FT-UNetFormer). Training parameters included:

(See Table 3 for detailed configuration)

The models were trained solely on RGB data, while the near-infrared band and nDSM were used exclusively for generating the reference data. The training involved augmenting tiles by horizontal and vertical flipping and added to the original patches, resulting in a total of 27,000 training patches. Training was stopped early if validation mIoU failed to improve for three consecutive epochs.

Model Performance and Accuracy Assessment

The performance was evaluated across 20 representative testing tiles, each with 5000×5000 pixels. The primary metrics used were mean Intersection-over-Union (mIoU) and mean F1 score (mF1) across all classes, emphasizing smaller TOF classes.

(See Table 5 for detailed class-wise metrics)

The FT-UNetFormer achieved the best overall results: mIoU score of 0.739 and mF1 score of 0.843. The model demonstrated high accuracy for the Forest class (IoU: 0.952, F1: 0.975), followed by TOF classes, with Linear showing the highest accuracy among them ("IoU:

Improvements for AI systems

As a fastidious and diligent AI researcher, I have analyzed this paper, Mapping and Classification of Trees Outside Forests using Deep Learning, which presents a systematic evaluation of state-of-the-art deep learning architectures for classifying four categories of woody vegetation (Forest, Patch, Linear, Tree) in agricultural landscapes.

The primary improvements to existing AI systems can be categorized into Architectural Advancements, Data Strategy Enhancements, and Model Deployment Capabilities.

Here are the specific improvements:


)Architectural Advancements

  1. [Vision Transformer Integration for Spatial Context] Implement or favor architectures like FT-UNetFormer over purely CNN-based models (e.g., U-Net). The paper demonstrates that vision transformers excel at capturing long-range spatial dependencies and context across distinct landscapes, which is critical for distinguishing complex TOF structures (Patch vs. Forest/Linear).

  2. [Hybrid Model Architectures] Develop robust hybrid models that combine the strengths of CNNs (for local feature extraction) and Transformers (for global context modeling). The success of FT-UNetFormer suggests that a dual-path approach is superior to standard encoder-decoder CNNs for this specific task.

  3. [Swin Transformer Backbone Utilization] Prioritize Swin Transformer backbones in future models due to their proven ability to capture multi-scale features efficiently, leading to superior performance (as seen with DC-Swin and FT-UNetFormer).

)Data Strategy Enhancements

  1. [Landscape-Specific Training Data Curation] Improve data strategy by explicitly curating training sets from regions with similar landscape structures (e.g., SH and NRW N) when developing models intended for deployment in novel agricultural environments, as the study showed that structural similarity drives generalization.

  2. [Class Imbalance Mitigation via Targeted Sampling] Implement advanced sampling techniques, specifically oversampling the underperforming or high-variability classes (like 'Patch') while undersampling dominant classes ('Forest'), to improve the model's ability to detect small woody features accurately.

)Model Deployment Capabilities

  1. [Robust Inference Pipeline for Large-Scale Mapping] Standardize the inference pipeline using a sliding window approach with overlapping predictions and majority voting, as demonstrated by FT-UNetFormer. This method ensures that boundary ambiguities are resolved by incorporating spatial context from neighboring predictions, leading to higher accuracy in final large-scale maps.

  2. [Automated Error Analysis Module] Integrate a mechanism to automatically generate and analyze normalized error matrices (as presented in Table 6) post-inference for any deployed model. This allows researchers to pinpoint specific failure modes (e.g., confusion between Linear and Background, or Patch vs. Forest) without requiring manual ground-truth verification across the entire dataset.

The improved AI system (specifically the FT-UNetFormer framework utilizing a Swin Transformer backbone) can perform the following:

  1. [High-Fidelity TOF Mapping]: Accurately delineate and classify trees outside forests into four distinct classes (Forest, Patch, Linear, Tree) with high mean Intersection-over-Union (mIoU) scores exceeding 0.74 across diverse agricultural landscapes.

  2. [Complex Structure Detection]: Specifically excel at detecting and segmenting the 'Patch' class—the most challenging category—achieving an mIoU of 0.606, which is significantly higher than baseline models in this difficult task.

  3. [Boundary Refinement]: Produce high-precision maps where boundaries between TOF classes (especially Linear) are well-defined, minimizing boundary confusion by leveraging overlapping predictions during inference.

  4. [Cross-Regional Transferability Assessment]: Provide a quantifiable metric for spatial generalization, allowing researchers to predict and assess the expected performance drop when deploying a model trained in one region to an entirely unseen agricultural landscape based on its structural similarity.

  5. [Policy Support Tool]: Serve as a scalable, automated tool for land-use planning and agroforestry policy by providing spatially comprehensive, fine-grained classification data that goes beyond simple binary TOF mapping (TOF vs. non-TOF).

Sources

Related papers