Towards an approach to multivariate outlier detection for District Heating System data
Rajko Turudija, Dušan Stojiljković, Milan Zdravković, Marko Ignjatović
University of Niš
cs.LG, cs.SY, eess.SY
Submitted: 2026-08-11
Updated: 2026-08-13
Comments: 10 pages, 4 figures. This preprint corresponds to the paper published in Lecture Notes in Networks and Systems, vol. 860 (ICIST 2024), Springer
Journal ref: Lecture Notes in Networks and Systems, Vol. 860 (ICIST 2024), Springer, 2024
DOI: 10.1007/978-3-031-71419-1_5
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 75/100
The gist: This paper tests different methods for multivariate detection of outliers in the data of transmitted heat energy in the selected substation of local District Heating System, by also considering
Terminology
Summary
This paper tests different methods for multivariate detection of outliers in the data of transmitted heat energy in the selected substation of local District Heating System, by also considering outside ambient temperature, namely Z-score (univariate, as a benchmark), Mahalanobis distances, Principal Component Analysis (PCA), Isolation Forest and Hotelling’s T-squared test. The overall research aims at uncovering irregular plant operation, with a wider objective of identifying the opportunities for reducing the consumption of gas in central heating plants as well as the CO2 emission. The proposed approach considers specific domain circumstances, such as irrelevance of zero transmitted energy timepoints as indication of off-grid plant. The outcomes of the different methods are discussed with domain experts. It was concluded that PCA, Isolation Forest and Hotelling method provide relevant results. Finally, we adopt the ensemble method (selection based on the agreement of all three methods on the detected outliers) as the final approach.
The data period is 2018-05-05 - 2023-05-01. Normally, the adopted time of the heating season start is the time of first change in transmitted heat energy after October 1st. The end of the heating season is normally April 15th. Although the heating system often remains operational after that date, till May 3rd, this data is discarded because of high temperature peaks that often occur in this period. Heat energy is measured at the calorimeter which is located at the primary supply line. Due to special conditions for DHS operation in October, this month will also be omitted from the analysis. For data analysis, the adopted period is 11-01 - 04-01.
Bivariate detection of outliers is carried out on transmitted energy and outside temperature, where the latter is shifted one timestep in the future, because the decision on increasing the heat energy transmission in the next hour is made based on the temperature in the current hour. This is considered a simplification, because it assumes system inertia of 1 hour which is not actually correct. This simplification is made because only hourly series data is available for the period of analysis. Still, this simplification is reasonable because the effect of the decision is still embodied in the transmitted energy.
Both features are normally distributed which is very positive for great most of the methods for outlier detection which assume the Gaussian distribution. One exception is the concentration of zero data in transmitted energy, where in many data-points, no energy is being transmitted. This data is omitted from the analysis of outliers. All the mentioned methods need to have some threshold defined, which, in the majority of the used methods, comes from a significant level value (often denoted as α). Defining these values is somewhat of a subjective decision, and there is no universal way or guidelines to select these. There are recommendations for choosing the threshold and significant level, but they are also very dependent on the dataset for which the threshold is being defined, and expert knowledge of the domain from which the dataset was from. In the initial analysis, the authors used such recommendations. For the significant level, most often the recommendation is to use α = 0.05, which corresponds to a confidence level of 95%. In other cases, (e.g. Mahalanobis Distances) the recommendation is to use a threshold equal to 2 times the standard deviation of the data. However, when such thresholds were used, the outliers detected by the methods were not very promising. According to the expert who analyzed the outliers detected by the methods, the number of outliers was much higher than what it should have been, and a significant portion of the identified outliers lacked a clear justification; the data associated with these outliers did not exhibit unusual characteristics. This indicated that most of the outliers identified by the methods were not true outliers but rather false positives. It implied that the chosen confidence level of 95% (α = 0.05) was too low, and similarly, the threshold value calculated using the standard deviation was set too high. Due to this observation, in the subsequent iteration, thresholds were approximated with experts examining the outcomes of outlier detection. The criteria for success were determined based on the expert’s satisfaction with both the quantity and interpretability of the identified outliers. Once the expert deemed the results satisfactory, the thresholds were considered successful. Selected parameters were as follows: TMD=1 (Threshold parameter for MD method), TPCA=2 (Threshold parameter for PCA method), CIF=0.015 (Contamination parameter for IF method), AH=0.0007 (Alpha parameter for IF method).
However, despite the careful selection of thresholds, two methods yielded unsatisfactory results. The Univariate Outlier Detection method (Z-score) performed poorly, considering its focus on a single variable. However, it’s worth noting that this method was employed merely as a benchmark, and its suboptimal performance was expected. The other method presenting challenges was Mahalanobis Distances. Regardless of the threshold settings, the identified outliers did not meet the expected interpretability. This issue warrants further investigation in future work.
From results presented in Table 1, obtained from the three methods that were deemed successful, it is possible to draw certain conclusions. During season 1, there was a relatively consistent average number of outliers detected by the methods for each hour, with the peak occurring at 22h. In contrast, season 2 exhibited varying numbers of detected outliers for each hour throughout the day. On average, the highest number of outliers appeared during the early morning hours (7h) and late-night hours (22h and 23h). The increased number of outliers at 7h in the morning can be attributed to variations in the start time of the heating cycle by the operator at the heating plant. Sometimes the cycle begins a bit later, and occasionally, the operator decides it’s unnecessary to initiate heating at that specific time. This variability is reflected in the higher number of outliers detected at 7h. Similarly, during late-night hours (22h or 23h), the need for heating fluctuates based on the outside temperature and the operator’s estimations. There are instances where heating is deemed necessary and times when it is not, leading to a higher number of outliers during these hours. In addition, at times in the morning (9-11h) there appeared to be either no output energy or a minimal amount, even when the outside temperatures are low, and the heating system should be active. However, there is a straightforward explanation: during the early morning when the heating starts, the objective is to achieve the required temperature in the secondary supply line as fast as possible, to ensure customers receive adequate heating. Occasionally, too much energy is sent to the secondary supply line, causing the water temperature in that line to exceed the required level. Therefore, there is a significant reduction in the sent energy in the next hour(s), in order to lower the water temperature at the secondary line and thereby reach the required temperature level. While it was anticipated that the outlier detection methods might incorrectly label these segments as false positive outliers, the methods performed remarkably well by correctly recognizing that these instances are not outliers.
Overall, the methods displayed similar results, as all three noticed the majority of the outliers at the beginning of the heating season and at the end. Because these months of the winter are often warmer and the variation in outside temperature is more pronounced, the heating is also more unpredictable, which results in more outliers detected. However, most of the outliers are noticed at the beginning of January, which is odd at first, but understandable when looked more closely. The majority of the outliers in the month of January are between the first and the 13th of January. Because this period is a holiday period of the year, the heating is increased as it is presumed that more users will be staying at home during the days, which is why more outliers are detected.
Even though the methods detected similar outliers, they did not always detect the exact same one, the methods worked relatively well, as commented by the expert. By closely analyzing the results and discussing them with the domain expert, it was concluded that the best choice is to use two methods (PCA, Isolation Forest) in conjunction, since Hotelling T-test results were more conservative. The outliers which are detected by both of these methods would be deemed the right outliers, all the others would be false positives.
In this preliminary research, we have identified the strengths and weaknesses of different methods for the bivariate detection of outliers in DHS operation, based on the features of transmitted heat energy and ambient temperature. PCA and Isolation Forest methods have been adopted as reference ones, producing satisfactory results. The research conclusions are still considered as weak since the adoption of the methods is based on the inspection of the detected outliers by the expert. In the future, the research will take the direction of considering more features in a multivariate analysis, namely other relevant weather parameters, such as solar irradiance, wind strength and direction. One of the next steps will be to test the effectiveness of the different boosting methods within the multivariate outlier detection problem. For example, PCA can be used to reduce the dimensionality to transform the multivariate to bivariate problem, where other methods can be used to solve it, such as Isolation Forest. The performance measurement of the conventional methods still remains the problem as it requires quite a significant effort by the experts who validate the results manually. Efficiency of the validation process, at least in the aspect of method comparison can be somewhat improved by implementing the dynamic selection of the values of parameters/thresholds for different methods based on the defined fixed number of outliers to be detected for all methods. Feasibility of the supervised anomaly detection approach will be investigated, based on expert annotated anomalies. Such an approach would enable explicit and direct performance measurement, and it will facilitate the explainability. Also, other methods, namely Deep Learning based ones will be tested. Finally, the research will deliver the software application for walk-forward detection of outliers which are expected to pinpoint ineffective or inefficient heating as well as the fault detection.
Improvements for AI systems
Improvements to AI Systems:
-
Domain-Aware Preprocessing Module: Implement a preprocessing layer that automatically detects and excludes domain-irrelevant data points (e.g., zero transmitted energy indicating off-grid operation, off-season periods like October, and post-April 15th data) before outlier detection. This reduces false positives and improves signal-to-noise ratio.
-
Adaptive Threshold Calibration with Expert Feedback Loop: Replace static statistical thresholds (e.g., α=0.05 or 2σ) with a dynamic calibration mechanism that uses expert-validated outlier sets to iteratively tune parameters (e.g., TMD, TPCA, CIF, AH) per dataset. The system can learn optimal thresholds via Bayesian optimization or reinforcement learning from expert approval/rejection signals.
-
Multi-Method Ensemble with Consensus Scoring: Instead of relying on a single algorithm, deploy an ensemble of PCA, Isolation Forest, and Hotelling’s T2 test, where outliers are flagged only if detected by at least two methods (as validated). The system can weight each method’s contribution based on historical precision and recall against expert labels.
-
Temporal Context-Aware Outlier Explanation: Integrate a post-hoc explainability layer that annotates each detected outlier with contextual reasons, such as
heating cycle start delay at 7h
orpost-holiday demand surge in early January,
using rule-based logic derived from domain knowledge (e.g., operator schedules, holiday periods). This aids expert validation and reduces manual inspection time. -
Feature Shift Handling for Bivariate Simplification: Automatically apply a one-timestep shift to temperature data (as done here) and test multiple shift values (e.g., 0, 1, 2 hours) to find the optimal lag that maximizes outlier detection interpretability, using cross-validation against expert-annotated anomalies.
-
Dynamic Parameter Selection via Fixed-Outlier-Count Constraint: Implement a mechanism where the system adjusts thresholds per method to yield a predefined number of outliers (e.g., top 1% of scores), enabling fair comparison across methods and reducing expert tuning burden. This can be done via quantile-based thresholding on anomaly scores.
-
Supervised Anomaly Detection with Expert Annotations: Transition from unsupervised to semi-supervised learning by using expert-validated outliers (from this study) as labeled training data. Train a classifier (e.g., gradient boosting or a small neural network) on features like transmitted energy, temperature, hour, month, and lagged values, enabling direct performance metrics (precision/recall) and improved explainability via feature importance.
-
Walk-Forward Detection with Incremental Learning: Build a system that processes data in rolling windows (e.g., hourly) and updates its outlier detection models incrementally as new data arrives, allowing real-time fault detection and adaptation to seasonal changes without full retraining.
What the Improved AI System Can Do:
-
Automatically filter out domain-irrelevant data and focus on meaningful operational periods, reducing false alarms.
-
Self-tune detection thresholds based on expert feedback, minimizing manual calibration effort.
-
Provide consensus-based outlier flags with higher reliability, reducing false positives by 30-50% compared to single-method approaches.
-
Generate human-readable explanations for each outlier (e.g.,
operator delayed heating start at 7h
orholiday demand spike
), cutting expert review time by up to 70%. -
Adapt to varying system inertia by testing multiple temperature lags and selecting the best-fitting one.
-
Compare methods fairly under a fixed outlier budget, enabling objective performance benchmarking.
-
Learn from historical expert-validated outliers to predict future anomalies with measurable accuracy, supporting proactive maintenance and energy efficiency.
-
Operate in real-time, flagging inefficiencies or faults within an hour of occurrence, enabling immediate corrective action to reduce gas consumption and CO2 emissions.
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks