Tabular Deep Learning vs Classical Machine Learning for Urban Land Cover Classification

arXiv:2609.19010 · cs.LG, cs.AI, cs.CV · Submitted 2026-07-20 · Read on arXiv

cs.LG, cs.AI, cs.CV

Submitted: 2026-07-20

Updated: 2026-07-20

Comments: Published in NeurIPS 2025 The 5th Muslims In ML (MusIML) Workshop

Code: https://github.com/mtesha/tdl-vs-ml-urbanlandcover

License: http://creativecommons.org/licenses/by/4.0/

The gist: Urban Land Cover (ULC) classification plays a crucial role in urban planning, environmental monitoring, and sustainable development.

Terminology

Abstract

Urban Land Cover (ULC) classification plays a crucial role in urban planning, environmental monitoring, and sustainable development. We study this task using the ULC dataset from the UCI Machine Learning Repository, which includes tabular features derived from high-resolution aerial imagery across nine classes (e.g., roads, trees, grass, water). The dataset presents typical remote sensing challenges, including high dimensionality, heterogeneous features, and class imbalance. In a unified, reproducible pipeline, we benchmark classical machine learning models (e.g., Logistic Regression, SVM, Random Forest, XGBoost, CatBoost) against Tabular Deep Learning (TDL) models (TabNet, FT-Transformer, TabTransformer, TabSeq, and 1D CNNs). To address class imbalance, we employ weighted cross-entropy loss for TDL models and evaluate performance using accuracy, macro-precision, macro-recall, macro-F1, AUC-ROC, and confusion matrices. Our results show that while tree ensembles remain strong general baselines, TDL models can match or exceed their performance when non-linear interactions are significant and imbalance handling is effective, providing complementary advantages for urban land cover mapping. See code: https://github.com/mtesha/tdl-vs-ml-urbanlandcover

Related papers