On the global feature importance for interpretable and trustworthy heat demand forecasting

arXiv:2608.13039 · cs.LG, cs.SY, eess.SY · Submitted 2026-08-13 · Read on arXiv

Milan Zdravković

University of Niš

cs.LG, cs.SY, eess.SY

Submitted: 2026-08-13

Updated: 2026-08-14

Comments: 9 pages, 5 figures. This preprint corresponds to the paper published in Thermal Science 2025 Volume 29, Issue 5 Part A, Pages: 3355-3365

Journal ref: Thermal Science 2025 Volume 29, Issue 5 Part A, Pages: 3355-3365

DOI: 10.2298/TSCI241223048Z

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 75/100

The gist: The paper introduces an ante-hoc Explainable AI (XAI) methodology to assess the global feature importance of Machine Learning models used for heat demand forecasting in intelligent control of

Terminology

Summary

The paper introduces an ante-hoc Explainable AI (XAI) methodology to assess the global feature importance of Machine Learning models used for heat demand forecasting in intelligent control of District Heating Systems (DHS), with the motivation to facilitate their interpretability and trustworthiness, hence addressing challenges related to adherence to communal standards, customer satisfaction, and liability risks. The methodology includes the use of four different approaches: intrinsic interpretability of the Gradient Boosting method and selected post-hoc methods, namely Partial Dependence (PD), Accumulated Local Effects (ALE), and SHAP. None of the selected methods assume feature permutation or perturbations which can introduce bias due to introduction of random unrealistic values of data instances. Discussion of results is provided, including the assessment of complementarities where applicable, with specific interpretations in context of the district heating processes.

The research aims at identifying effective approaches to evaluating trustworthiness of heat demand forecasting models by using different aspects of interpretability of feature importances at the global level, for the whole trained model. The test model is trained using ensembles of decision trees (XGBoost algorithm) to make hourly forecasts of transmitted heat energy using 4 heating seasons of historical data. Historical data from SCADA system of substation 17 of local DHS, managed by the Faculty of Mechanical Engineering, is used, encompassing 4 heating seasons (2020-2024) and merged with meteorological data. Data includes hourly transmitted heating energy (deltae), ambient temperature (tamb), water temperatures in secondary supply and return lines (tsups ec, trets ec), water temperatures in primary supply and return lines (tsupp rim, tretp rim), and various meteorological data. Preprocessing includes stripping data points, inserting missing timepoints, treatment of zero data, introducing time features, linear interpolation, removing data outside heating season, removing features with Pearson correlation coefficient less than 0.1, replacing small energy values with zero, and removing outliers (z-score>4). Data is enriched with time-lagged transmitted energy (deltae-1 to deltae-23) and ambient temperature (temp-1 to temp-23) in the past 24 hours. 80% of data is used for training and 20% for testing, with Mean Absolute Error (MAE) as the metric.

Global feature importance is discussed using intrinsic interpretability of gradient boosting and selected ante-hoc methods. Partial Dependence (PD) measures how one model prediction varies with respect to a feature of interest, with the main limitation of assuming feature independence. Accumulated Local Effects (ALE) plots are a natural extension that doesn't struggle with dependencies in underlying features. SHAP (SHapley Additive exPlanations) values come from cooperative game theory and are used to fairly distribute total gain among features based on their contributions, and can be used for both local and global explainability.

For the Gradient Boosting method, the paper discusses gain, cover, and frequency metrics. Gain measures improvement in accuracy that a feature provides when used in a split. Cover measures the relative quantity of observations affected by a feature when used in a split. Frequency (or weight) measures the number of times a feature is used in all trees. The weight bar plot illustrates that ambient temperature (tamb) is the most used feature in decision making across multiple trees, indicating it is an important decision point. Transmitted energy at the same hour of the previous day (deltae-23) is by far the most influential in improving model performance when used in splits, indicating strong daily seasonality of DHS operation. Temperature reading by the sensor near the substation in the same hour of the previous day (temp-23) has the larger cover, meaning it was a relevant decision point for most instances. These interpretations correspond to common sense and real-life situations, meaning the trained model accurately reflects complex dependencies within DHS.

For Partial Dependence, ICE plots are plotted for four selected features: deltae-23, hour of the day, tamb, and temp. The PD line indicates that higher values of tamb are generally associated with a decrease in predicted value, but this trend reverses at approximately tamb=13, indicating non-linear effects or feature interactions. The PD plot of deltae-23 is mostly flat, especially for delta-23<300. The PD plot of hour of day follows the change pattern from the time series, confirming repeatability of this pattern and strong hourly seasonality, with predictions significantly higher during certain hours (early morning to late afternoon) and dropping sharply in late evening and early night hours. The flatness of deltae-23 PD plot could imply its non-existing effect on target variable despite high weight, but ICE plots show that the average is not representative of contributions over individual instances having quite diverse plots.

For Accumulated Local Effects, the distribution of ALE for temperatures is more stable than in the case of PDP. It shows that transmitted energy will decrease almost linearly in the interval of tamb<9, then the decrease rate will become more rapid with some non-linearities. Flatness of ALE plot for deltae-23 feature is explained as previously.

For SHAP, the global SHAP feature importance is determined by averaging absolute SHAP values for each feature across all instances. The SHAP summary plot shows that features with the widest range of SHAP values, namely deltae (transmitted energy at current timepoint) and hour of day, have the greatest impact on model output. High values of deltae increase the forecast (SHAP values > 0), while low values decrease it. It's clearly the opposite case for trets ec and feelslike. Notable differences with gain, cover, and weight metrics are clarified: intrinsic explainability features are rooted in training data, while post-hoc interpretability feature importances are calculated based on test data. Weight and cover are complementary to SHAP importances and cannot be compared with it. Comparing gain with SHAP values can validate whether features that improve model accuracy most are also those that most influence predictions. In DHS forecasting model, tamb has moderately high SHAP values but is not ranked in the gain bar plot, suggesting redundancy. deltae-23 has high gain but low SHAP values, possibly because its effect is overshadowed by other features or only relevant in specific scenarios, such as the beginning of the workday when heating starts.

The paper concludes that this research introduces global feature importance calculation methods with intended use to validate trained forecasting models in environments requiring legal compliance and agreed service levels on a massive customer scale, which is all the case with District Heating services. Four methods are implemented: ante-hoc feature importances, Partial Dependence, Accumulated Local Effects, and SHAP. Some widely used techniques (Permutation Importance and LIME) were discarded because they assume feature permutation or perturbations introducing bias. Both PD and ALE use marginal distributions but only with minor variances. Partial Dependence captures only main effects and ignores feature interactions, but ICE plots allow detection of heterogeneous effects. ALE has been proven as more reliable method, taking into account feature interactions and nonlinearities. ALE plots clearly reveal that transmitted energy decreases almost linearly in the interval of tamb<9, with more rapid decrease rate for warmer temperatures, closely resembling the shape of the control curve, which is strong evidence of model quality. SHAP's assessment is based on game theory, accounts for feature interactions and nonlinearities, making it the most reliable and robust method. Benefits of using XAI techniques include transparency helping build trust, identifying key features guiding operational strategies, efficiency improvements, supporting long-term planning, and refining the model.

Improvements for AI systems

Improvements to AI Systems:

  1. Implement Ante-Hoc Global Interpretability for Compliance-Critical Forecasting: Integrate the described multi-method XAI framework (intrinsic Gradient Boosting metrics + PD/ALE/SHAP) directly into the model validation pipeline. This allows the AI to automatically generate compliance reports for regulatory bodies, demonstrating that feature importance aligns with physical domain knowledge (e.g., ambient temperature and time-lagged energy) without relying on biased permutation-based methods.

  2. Enhance Robustness by Avoiding Perturbation-Based Explanations: Replace Permutation Importance and LIME with the paper’s chosen methods (PD, ALE, SHAP) to prevent the AI from generating misleading explanations due to random unrealistic feature values. The improved system will produce stable, trustworthy interpretations even when features are highly correlated (e.g., temperature and lagged energy), reducing liability risks in customer-facing decisions.

  3. Enable Detection of Non-Linear and Interaction Effects: Use ALE plots (instead of PD alone) to capture feature interactions and non-linearities, such as the piecewise linear decrease in heat demand with ambient temperature below 9°C. The AI can then identify threshold effects and adapt control strategies accordingly, improving operational efficiency in District Heating Systems.

  4. Resolve Contradictory Feature Importance Signals: The system will cross-validate intrinsic metrics (gain, cover, weight) with post-hoc SHAP values to distinguish between features that improve model accuracy (e.g., deltae-23) and those that drive actual predictions (e.g., deltae, hour of day). This allows the AI to flag redundant features (like tamb with high SHAP but low gain) and refine feature engineering, reducing model complexity without losing predictive power.

  5. Improve Temporal Pattern Recognition for Load Forecasting: By leveraging the paper’s findings on strong daily and hourly seasonality (e.g., high SHAP for hour of day, high gain for deltae-23), the AI can prioritize time-lagged features and time-of-day encoding in its architecture. This leads to more accurate hourly heat demand forecasts, especially during workday startup periods, enabling proactive control of substations.

  6. Provide Instance-Level Heterogeneity Insights: Incorporate ICE plots alongside PD to detect when average effects mask diverse individual behaviors. The AI can then identify specific scenarios (e.g., early morning hours) where the model’s predictions are less stable, triggering targeted retraining or manual oversight to improve reliability.

  7. Support Long-Term Planning and Scenario Analysis: Use ALE and SHAP to simulate how changes in external conditions (e.g., warmer winters) affect heat demand. The improved AI can generate what-if analyses for infrastructure investment, tariff setting, and carbon reduction strategies, based on validated feature-response curves rather than black-box outputs.

  8. Automate Model Quality Assurance with Physical Consistency Checks: The system will compare XAI outputs against known DHS control curves (e.g., the temperature-response shape) as a sanity check. If explanations deviate from expected physical behavior, the AI can automatically flag data quality issues or model drift, ensuring continuous trustworthiness in production.

Sources

Related papers