Remote Sensing and Machine Learning-Based Analysis of Land Use and Vegetation Change in Dhaka District, Bangladesh

arXiv:2608.12001 · cs.LG, cs.AI, cs.CV · Submitted 2026-08-12 · Read on arXiv

Muhammad Masud Tarek, Md. Alamgir Hossain, Md. Samiul Islam, Muntasir Hasan Kanchan

State University of Bangladesh · American International University - Bangladesh · Skill Morph Research Lab · Skill Morph

cs.LG, cs.AI, cs.CV

Submitted: 2026-08-12

Updated: 2026-08-13

Comments: 16 Pages, 8 Figures, 9 Tables

License: http://creativecommons.org/licenses/by-nc-nd/4.0/

Importance score: 51/100

The gist: This study employs remote sensing data and machine learning techniques to analyze spatiotemporal changes in land cover and vegetation dynamics in Dhaka District, Bangladesh, between 2019 and 2024.

Terminology

Summary

This study employs remote sensing data and machine learning techniques to analyze spatiotemporal changes in land cover and vegetation dynamics in Dhaka District, Bangladesh, between 2019 and 2024. High-resolution satellite imagery from Sentinel-2 MSI and Landsat 8 was utilized to classify land cover types and compute spectral indices including the Normalized Difference Vegetation Index (NDVI), Normalized Difference Built-up Index (NDBI), and Normalized Difference Water Index (NDWI). A supervised machine learning approach incorporating Decision Tree, K-Nearest Neighbors (KNN), and Random Forest classifiers was applied using labeled geospatial training points within Google Earth Engine. Accuracy assessments were conducted using confusion matrices and kappa statistics. Results indicate a 59.5% increase in urban built-up areas and a significant decline in vegetation (−8.46%) and water bodies (−7.77%) over the five-year period. Land conversion from vegetated and aquatic areas to urban infrastructure was identified as a dominant trend. Among the models, Random Forest demonstrated the highest classification accuracy. These findings underscore the growing environmental pressures driven by unregulated urban expansion in Dhaka. The study highlights the potential of remote sensing and machine learning tools in providing timely, actionable data to support sustainable urban development, land-use regulation, and ecosystem conservation policies.

The study area, Dhaka District, is characterized by rapid urbanization and dense population, encompassing both urban and peri-urban areas with diverse land cover types including residential, commercial, agricultural, and water bodies. The research utilized Sentinel-2 MSI multispectral imagery, with administrative boundary data collected from the Food and Agriculture Organization of the United Nations (FAOUN) GAUL 500m dataset. Preprocessing steps included radiometric calibration, geometric correction, cloud masking, and image enhancement. Cloud removal was performed using cloud masking and filtering techniques, temporal compositing, and multi-date image mosaicking, with images having less than 30% cloud coverage used to generate cloud-free composites.

For supervised classification, 430 labeled geospatial points were manually selected and categorized into four land cover classes: water body (101 points), vegetation (153 points), built-up/urban area (133 points), and soil/bare land (43 points). The ESA WorldCover 10m 2021 pretrained dataset was used as a training reference and benchmark. The ESA WorldCover classification for 2021 showed cropland as the most prominent class covering approximately 632 square kilometers, followed by tree cover (413 km2), built-up areas (239 km2), bare/sparse vegetation (75 km2), grassland (22 km2), permanent water bodies (79 km2), and herbaceous wetland (7 km2), totaling 1,467 square kilometers, remarkably close to the official administrative area of 1,471 km2.

The performance comparison for the 2019 classification showed that Random Forest consistently outperformed other models, achieving the highest F1-scores across most land classes and the strongest kappa coefficient of 0.8362. Water and vegetation classes were classified with very high accuracy in all models, while bare land presented challenges with lower F1-score values. The Random Forest model achieved an overall accuracy of 88.35% for both 2019 and 2024 classifications.

Quantitative comparison of land cover areas between 2019 and 2024 revealed a 53 square kilometer increase in urban area (from 89 km2 to 142 km2), signaling rapid urban expansion. Vegetation decreased by 22 km2 (from 969 km2 to 947 km2), water bodies shrank by 46 km2 (from 244 km2 to 198 km2), and bare land experienced a 15 km2 increase (from 165 km2 to 180 km2). The interclass transition matrix showed that 62 km2 of land classified as vegetation in 2019 was converted into urban area by 2024, while an additional 12 km2 of bare land also transitioned to urban use.

The spectral index-based analysis using NDVI, NDWI, and NDBI revealed that vegetation declined from 945 km2 in 2019 to 865 km2 in 2024—a net loss of 80 km2, representing an 8.46% decrease. Urban area expanded from 278 km2 to 302 km2—an increase of 24 km2, approximately 8.63% growth. Water bodies reduced from 90 km2 to 83 km2, marking a 7.77% decline. The study noted methodological differences between index-based and machine learning approaches, with index-based methods focusing only on three land classes (vegetation, water, urban), excluding bare land, and depending heavily on threshold values.

The research answers revealed that urban areas expanded primarily in the north and central parts of Dhaka, often replacing previously vegetated zones, while vegetation loss was prominent across peripheral and agricultural zones, and water body reduction was observed in the southern and southwestern regions. The findings emphasize the urgency of implementing sustainable urban planning and environmental conservation strategies, particularly in rapidly urbanizing regions like Dhaka. Future research can explore the integration of higher-resolution commercial satellite imagery and deep learning models such as convolutional neural networks (CNNs) to improve classification accuracy and detect finer-scale urban dynamics, as well as incorporating socio-economic and climate data to enhance understanding of drivers behind land cover change.

Improvements for AI systems

Improvements to AI Systems:

  1. Hybrid Classification Pipeline: Integrate Random Forest as the primary classifier but add a secondary refinement layer using a rule-based post-processing step that leverages NDVI/NDBI/NDWI thresholds to correct misclassifications (e.g., reclassify bare land mislabeled as built-up when NDBI is low). This combines the high accuracy of machine learning with the interpretability of spectral indices.

  2. Temporal Change-Aware Training: Train the model on multi-temporal composites (e.g., seasonal mosaics from 2019 and 2024) rather than single-date images. This allows the AI to learn phenological variations, reducing false positives in vegetation/water classification due to seasonal changes (e.g., monsoon vs. dry season).

  3. Uncertainty-Aware Segmentation: Implement a probabilistic output layer (e.g., softmax probabilities from Random Forest) to flag low-confidence pixels. The improved system can then automatically trigger a deep learning model (e.g., U-Net) for those specific regions, improving accuracy in heterogeneous urban-peri-urban interfaces where bare land and built-up areas overlap.

  4. Automated Threshold Optimization: Replace manual threshold selection for NDVI/NDWI/NDBI with an AI-driven optimizer (e.g., Bayesian search) that learns optimal cutoffs per land cover class and per season. This reduces the 8.46% vegetation loss overestimation caused by fixed thresholds, as noted in the paper’s methodological limitations.

  5. Transition Matrix Prediction Module: Add a recurrent neural network (LSTM) trained on the 2019–2024 interclass transition matrix to predict future land cover changes (e.g., 2029). The improved system can forecast urban expansion hotspots, enabling proactive land-use regulation rather than reactive monitoring.

  6. Cross-Sensor Domain Adaptation: Build a domain-adversarial neural network that aligns Sentinel-2 and Landsat 8 feature distributions. This allows the AI to fuse both sensors seamlessly, improving temporal consistency and reducing sensor-specific biases in change detection (e.g., the 53 km2 urban increase vs. 24 km2 index-based discrepancy).

What the Improved AI System Can Do:

  • Produce high-resolution, temporally consistent land cover maps with >90% overall accuracy by combining Random Forest with uncertainty-guided deep learning refinement.

  • Automatically generate early-warning alerts for vegetation loss or water body shrinkage in specific administrative zones (e.g., southern Dhaka) by comparing predicted 2029 maps against sustainability targets.

  • Quantify land conversion drivers (e.g., agricultural-to-urban) with per-pixel attribution, linking each change to specific spectral signatures, enabling policymakers to target regulation at the most affected parcels.

  • Adapt to new geographic regions without manual retraining by using the pretrained ESA WorldCover reference and transfer learning, making it deployable for other rapidly urbanizing South Asian cities.

  • Deliver interactive scenario simulations (e.g., if urban growth continues at 59.5% per 5 years, water bodies will shrink by X km2 by 2034) by coupling the LSTM predictor with the transition matrix, supporting evidence-based urban planning.

Abstract

Rapid urbanization in Dhaka District, Bangladesh has triggered substantial alterations in land use and environmental conditions, necessitating systematic monitoring for informed urban planning and ecological sustainability. This study employs remote sensing data and machine learning techniques to analyze spatiotemporal changes in land cover and vegetation dynamics between 2019 and 2024. High-resolution satellite imagery from Sentinel-2 MSI and Landsat 8 was utilized to classify land cover types and compute spectral indices including the Normalized Difference Vegetation Index (NDVI), Normalized Difference Built-up Index (NDBI), and Normalized Difference Water Index (NDWI). A supervised machine learning approach incorporating Decision Tree, K-Nearest Neighbors (KNN), and Random Forest classifiers was applied using labeled geospatial training points within Google Earth Engine. Accuracy assessments were conducted using confusion matrices and kappa statistics. Results indicate a 59.5% increase in urban built-up areas and a significant decline in vegetation (-8.46%) and water bodies (-7.77%) over the five-year period. Land conversion from vegetated and aquatic areas to urban infrastructure was identified as a dominant trend. Among the models, Random Forest demonstrated the highest classification accuracy. These findings underscore the growing environmental pressures driven by unregulated urban expansion in Dhaka. The study highlights the potential of remote sensing and machine learning tools in providing timely, actionable data to support sustainable urban development, land-use regulation, and ecosystem conservation policies.

Related papers