XGBoost "is all you need": the case of forecasting transmitted heat energy in District Heating Systems

arXiv:2608.11446 · cs.LG, cs.SY, eess.SY · Submitted 2026-08-11 · Read on arXiv

Milan Zdravković

University of Niš

cs.LG, cs.SY, eess.SY

Submitted: 2026-08-11

Updated: 2026-08-13

Comments: 9 pages, 7 figures. This preprint corresponds to the paper published in Lecture Notes in Networks and Systems, vol. 860 (ICIST 2024), Springer

Journal ref: Lecture Notes in Networks and Systems, Vol. 860 (ICIST 2024), Springer, 2024

DOI: 10.1007/978-3-031-71419-1_2

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 65/100

The gist: This paper presents a comparative study of two distinct approaches, XGBoost and Long-Short Term Memory (LSTM), for forecasting transmitted heat energy in District Heating Systems (DHS).

Terminology

Summary

This paper presents a comparative study of two distinct approaches, XGBoost and Long-Short Term Memory (LSTM), for forecasting transmitted heat energy in District Heating Systems (DHS). The objective is to explore scenarios in which conventional ML algorithms demonstrate better performance over deep learning networks in time series forecasting and the associated benefits in terms of computational cost and environmental impact. The study focuses on a real-world DHS dataset. Through experimentation and analysis, it is demonstrated that XGBoost consistently outperforms LSTM in this specific forecasting task. The difference is explained by the error distribution illustrating that LSTM makes more significant errors in the intervals of less data availability. The reduced computational demands of conventional ML approaches not only result in cost savings but also minimize the carbon footprint associated with data analysis tasks in energy systems.

The paper argues that the extreme popularity of off-the-shelf deep learning algorithms has led to hasty adoption without careful consideration of algorithm suitability, data quality, and computational requirements, often resulting in systems that underperform or consume excessive energy. The research objective is to demonstrate that traditional ML algorithms are competitive compared to complex neural networks in certain time series forecasting problems, showcased on forecasting transmitted heat energy in DHS.

The methodology involves comparing a stacked LSTM architecture and XGBoost for forecasting transmitted heat energy based on measured air temperature from the previous timepoint. Data from one substation (substation 9) from a local DHS was used, including outside ambient temperature and transmitted energy from two heating seasons (2018/19 and 2019/20), totaling 5832 timepoints. Only the period from November to March was considered. Missing data due to lack of 3G network connectivity was imputed using linear interpolation. Both time series were stationary, confirmed by the Augmented Dickey-Fuller test. The distribution analysis indicated relatively high sparsity of hourly transmitted energy, where zeros indicate the system is not operational, mostly in cases of high outside temperatures. The Spearman coefficient (-0.308) indicated a statistically significant negative association between the two signals.

For the LSTM model, a stacked architecture with two LSTM layers each with 100 units and ReLU activation, each followed by a dropout layer, was used. Data was normalized, and training was carried out with 100 epochs and batch size of 24, with 20% of the training set used for validation. For the XGBoost model, time series data was transformed into structured format using lagged features (6 timepoints in the past) and time-based features (hour of day, day of week, month). Spearman correlation showed statistically significant association between hour of day and month with transmitted heat energy (SPhour=-0.339, SPmon=-0.150), but not for day of week. Bayesian optimization using Hyperopt with Tree of Parzen Estimators was used for hyperparameter optimization over 80 iterations, with MAE as the objective function.

The results showed XGBoost outperformed LSTM in all metrics. The optimized XGBoost achieved RMSE of 35.677 kW, MAE of 18.687 kW, and R2 of 0.861, compared to LSTM's RMSE of 62.111 kW, MAE of 28.890 kW, and R2 of 0.540. Training time for LSTM was 476.129 seconds, while XGBoost was 0.238 seconds (549.237 seconds with optimization). Inference time was 0.537 seconds for LSTM versus 0.007 seconds for XGBoost (0.035 seconds optimized). Training CO2 emissions were 9.245 g for LSTM versus 0.005 g for XGBoost (10.664 g optimized), and inference CO2 emissions were 10.421 g/1000 for LSTM versus 0.141 g/1000 for XGBoost (0.676 g/1000 optimized).

The error distribution histograms showed that while both models had good accuracy with few substantial errors, LSTM had minor fat tails with increased accumulation of larger errors, explaining the difference in performance metrics. Scatter plots of actual versus predicted values showed points clustered along the line but with more variance for LSTM, especially for lower and higher values where data availability is limited. Both models struggled to forecast periods when heating is turned off by the operator, which is a human decision not included as a feature.

The paper identifies several reasons for XGBoost's superior performance: the small dataset size combined with feature engineering that encapsulates temporal dynamics well, LSTM's sensitivity to hyperparameter choices, and the fact that LSTM requires data imputation which introduces bias, while XGBoost can handle missing data naturally and can use data islands occurring during sensor faults. Additionally, XGBoost offers better interpretability through feature importances, which showed the highest importance of transmitted energy in the current hour, current ambient temperature, and hour of the day.

The conclusion states that conventional ML approaches should be the first choice for relatively smaller datasets and datasets with sparse or missing data, as they are not only more sustainable in terms of computational requirements but also more accurate. The necessary condition for this better performance is feature engineering practices that reflect domain experience and expert knowledge. Standard ML algorithms also offer much better interpretability, which is crucial in industrial applications.

Improvements for AI systems

Improvements to AI Systems:

  1. Implement a model-selection framework that automatically evaluates traditional ML algorithms (e.g., XGBoost, Random Forest) against deep learning models (e.g., LSTM, Transformer) before deployment, using dataset size, data sparsity, and missing-data patterns as decision criteria. This prevents hasty adoption of complex neural networks for small or sparse datasets.

  2. Add a missing-data-aware preprocessing module that detects data islands and gaps, then either imputes with linear interpolation (for LSTM) or uses native missing-value handling (for XGBoost) based on which yields lower validation error. The system can dynamically choose the best strategy per feature.

  3. Integrate a feature-engineering layer that automatically generates lagged features (e.g., 6 previous timepoints) and time-based features (hour, day, month) for structured models, while also providing the raw sequence for recurrent models. This dual-format input allows fair comparison and enables the system to select the best representation.

  4. Embed a computational-cost and carbon-footprint estimator into the training pipeline, reporting training time, inference time, and CO2 emissions (e.g., g CO2 per run) alongside accuracy metrics. The system can then recommend the most sustainable model that meets a user-defined accuracy threshold.

  5. Add an error-distribution analysis tool that compares prediction errors across data-density intervals (e.g., low, medium, high availability). This helps identify if a model fails specifically in sparse regions, triggering a switch to a more robust algorithm or additional data augmentation.

  6. Implement an interpretability module for any chosen model, using feature importance (for XGBoost) or SHAP values (for LSTM) to explain predictions. This is critical for industrial adoption, allowing operators to trust and audit the forecasting system.

  7. Create an adaptive retraining trigger that monitors real-time data distribution shifts (e.g., seasonal changes, sensor faults) and automatically retrains or switches between XGBoost and LSTM based on recent validation performance, ensuring sustained accuracy without manual intervention.

What the Improved AI System Can Do:

  • Automatically recommend and deploy the most accurate, cost-effective, and environmentally friendly forecasting model for a given time-series dataset, avoiding unnecessary deep learning complexity.

  • Handle missing or sparse data intelligently by choosing between imputation and native missing-value handling per model type.

  • Provide transparent, interpretable forecasts with feature importance rankings, enabling domain experts to validate and adjust predictions.

  • Operate with minimal computational resources (e.g., 0.238s training vs. 476s for LSTM) and near-zero carbon footprint for small-to-medium datasets, making it deployable on edge devices or in real-time industrial settings.

  • Continuously self-optimize by comparing error distributions and switching algorithms when data conditions change, ensuring robust performance across heating seasons, sensor outages, or operational changes.

Related papers