Air Quality Station Simulation via LSTM and Attention-Based Modelling
Alexander Kostadinov, Petar O. Hristov, Dessislava Petrova-Antonova
GATE Institute · Sofia University St Kliment Ohridski
cs.LG
Submitted: 2026-08-12
Updated: 2026-08-13
Comments: Preprint prepared as an extension to the conference article at https://doi.org/10.1007/978-3-031-97313-0_20, which describes some features of the proposed model. The preprint includes more tests and better concept explanations
Project page: https://www.who.int/news-room/fact-sheets/detail/ambient-
License: http://creativecommons.org/licenses/by-sa/4.0/
Importance score: 75/100
The gist: The paper presents SATADL (SpAtial-Temporal Attention Dual LSTM), a deep-learning model designed to simulate the measurements of an unresponsive air quality station until its operation is restored.
Terminology
Summary
The paper presents SATADL (SpAtial-Temporal Attention Dual LSTM), a deep-learning model designed to simulate the measurements of an unresponsive air quality station until its operation is restored. The model can infer complex relations and output multiple-hour-ahead air-quality forecasts.
The architecture allows it to extract information from different aspects of the data
through spatial and temporal modules with attention mechanisms. The model was demonstrated on four sets of air quality stations from around the world (Beijing, Hong Kong, Sofia, and Delhi), simulating PM10 concentrations for periods of hypothetical station failures lasting up to 48 hours. Results show that SATADL performs better across different prediction windows, for both coefficient of determination and root mean squared error, demonstrating its suitability as a virtual proxy station.
The paper addresses the problem that poor air quality in urban areas is driven by a complex chain of processes and presents a significant public health concern.
Air pollution leads to multiple diseases and complications
with an estimation of 6.7 million deaths globally, according to the World Health Organisation (WHO).
Air quality is used as a natural progress indicator towards the transformation of urban areas into smart cities,
with urban digital twins serving as digital replicas of different dimensions and processes of the city.
Municipalities have established monitoring systems measuring pollutants such as particulate matter no larger than 10 µm in aerodynamic diameter (PM10), nitrogen dioxide (NO2), sulphur dioxide (SO2)
and meteorological factors. However, certified air quality stations are expensive equipment and are usually distributed at only a few carefully chosen locations, making each individual station critical to understanding the quality of air in the city.
These systems are susceptible to technical issues, weather events, and power outages, that compromise their reliability as measurement devices.
The paper notes that the development of models for real-time simulation of air quality stations during extended periods of hardware downtime, has not been addressed in the literature, to the best of the authors' knowledge.
SATADL is an encoder-decoder deep neural network
with the following modules:
The spatial module processes the input from the surrounding stations and dynamically captures their impact on the output.
It comprises:
-
Time Feature Embedding:
Each time feature, wi, in the input sequence is embedded using three separate embedding layers, corresponding to hi, di and mi
(hour, day, month), with positional encoding using sine and cosine transformations for periodicity. -
Spatial Attention:
The spatial attention mechanism captures the dependency between the functioning stations and their influence on the predictions.
It uses additive attention toassign dynamic weights, according to the impact each station has on the prediction.
The previous hidden and cell states from the LSTM block are fed back in, and GELU activation is applied after linear layers.
The temporal module captures the influence of past target-station measurements on future predictions.
It includes:
- Temporal Attention:
The goal of the temporal attention is to highlight significant moments in the past data from the target station and form a temporal context by assigning weighted scores to each time step.
It usesscaled dot-product attention
to extract the temporal context, forming query, key, and value matrices from the expanded measurements concatenated with time feature embeddings.
The decoder combines a uni-directional LSTM (uLSTM) block, which extracts local temporal patterns, and a bi-directional LSTM (biLSTM) block, which processes spatial contexts to produce base values.
The uLSTM produces a representation of the local patterns, which serve as a trend adjustment.
The biLSTM works sequentially by adding each of the spatial contexts
and produces base predicted values that are concatenated and summed with the output of the uLSTM block to form the final prediction.
The problem involves a malfunctioned target
station that needs simulation using data from its historical measurements and present data from the other functioning stations.
Time steps are divided into past (Tp, with P steps) and present (Tf, with F steps). The model uses:
-
F: data from functioning stations at present time steps
-
P: past data from the target station
-
W F and W P: time features (hour, day, month) for present and past steps
Four datasets were used:
-
Hong Kong: 16 stations, measuring SO2, NO2, NOx, CO, O3, PM2.5, PM10, with meteorological conditions T, WS, WD, RH; period 2016-2019; target station
Central
-
Delhi: 9 stations, measuring NO, NO2, NOx, CO, O3, PM2.5, PM10, with T, P, R, DPT, WS; period 2019-2024; target station
Pusa
-
Sofia: 5 stations, measuring SO2, NO, NO2, PM10; period 2015-2024; target station
Nadezhda
-
Beijing: 13 stations, measuring SO2, NO2, CO, O3, PM2.5, PM10; period 2013-2017; target station
Huairou
The data contained missing and negative values, as well as outliers.
The cleaning steps included:
-
Negative values, except those for temperature, were treated as missing
-
Longer periods of repeating values were taken to indicate anomalous readings
using rolling-window standard deviation thresholds -
Large gaps in the data were removed
with intervals longer than four hours classified as large -
Linear interpolation was applied
for remaining missing points -
Data was filtered to retain only timestamps where measurements are available from all stations
The augmented Dickey-Fuller test
showed no evidence suggesting statistically significant non-stationarity in any of the attributes in all data sets.
-
Adam optimizer with learning rate 0.001 and reduction on plateau of 0.5
-
Loss function: mean squared error
-
Prediction steps F = 12 hours and F = 48 hours
-
Past steps P = 12 hours
-
30 epochs
-
Data divided into training, validation, and test sets with ratio 8:1:1
-
PM10 chosen as target pollutant
because it is measured across all station networks and carries regulatory significance
SATADL was compared against five models:
-
LSTM:
A single layer unidirectional long short-term memory neural network with a projection layer
-
Attention-LSTM:
A model proposed by Gangopadhyay et al. (2018), which uses stacked LSTM layers and attention mechanism
-
Transformer:
A sequence-to-sequence model developed by Vaswani et al. (2017)
-
CLR:
A combined model of convolution, fully connected and recurrent layers
-
LSTNet:
A model proposed by Lai et al. (2018) combining convolution and recurrent layers
Spatio-temporal variants of each benchmark model were developed, comprising two copies of the respective model, one in a spatial and one in a temporal block (similar to SATADL).
Most models are able to deal with short-term faults up to twelve hours with high R2 and low RMSE. In all datasets SATADL achieves the best performance.
Key findings:
-
The transformer exhibits the weakest performance of all models
with R2 = 0.17 on Delhi data -
The attention-LSTM showed strong results across all datasets
-
The CLR and LSTNet models... show mixed results
-
SATADL
shows superior performance in capturing the qualitative nature of the observed data
None of the models performs well over the Delhi data, despite SATADL ranking highest in both metrics.
The paper attributes this to the impact of local patterns in the data, coupled with weak correlations between the target and surrounding stations.
SATADL achieves substantially better average R2 scores in all forty-eight-hour scenarios
and widens its lead in the Delhi and Sofia datasets.
Analysis at hours 1, 6, 12, 24, 36, and 48 showed:
-
All models gradually lose performance as the time progresses
-
SATADL combines the benefits of high accuracy in the early stages with low error propagation, which results in the best performance at each time step
-
"Because of the combination of information from the functioning stations at each hour and the past data and its trend, SATADL can capture the global levels of the target pollutant, while preserving the local patterns detected at the target station"
-
The attention in the spatial module automatically removes any noisy or weakly-contributing station measurements by giving them very low weights
The paper identifies two main limitations:
-
SATADL assumes that all surrounding stations operate while the target is not functioning
-
The model can only work with time series data
- static data such as building and infrastructure could improve results but would reduce usability
Future directions include:
-
Integrating SATADL into a real-time air quality data collection pipeline
-
Comparison against bigger transformer-based models
-
Development of a more adaptive solution, that deals with time periods when multiple stations are non-operational
-
Development of a new module that can accept static data
-
Experiment with multiple pollutant predictions
-
Application
into scenarios more abstractly resembling physical sensor networks, for example, in increasing the reliability of heavily instrumented critical systems
The paper concludes that SATADL learns local and global patterns better, resulting in more stable and realistic behaviour of the simulated measurements and achieves better R2 and RMSE results in both short- and long-term failure scenarios.
The model is described as a robust simulation model designed to address various scenarios involving data unavailability and sensor malfunctions in scenarios different to air quality modelling.
Improvements for AI systems
Improvements to AI Systems Based on SATADL
- Hybrid Attention Fusion for Multi-Source Sensor Imputation
-
Improvement: Integrate SATADL’s dual-module design (spatial attention + temporal attention) into a general-purpose missing-data imputation framework. Replace the LSTM backbones with lightweight temporal convolutional networks (TCNs) or state-space models (e.g., Mamba) to reduce inference latency while retaining the additive and scaled dot-product attention mechanisms.
-
Capability: The improved system can impute missing values across heterogeneous sensor networks (e.g., traffic, weather, energy grids) with dynamic weighting of correlated sources, automatically down-weighting noisy or irrelevant sensors, and preserving local temporal trends over long horizons (up to 48+ steps).
- Adaptive Multi-Station Failure Handling
-
Improvement: Extend SATADL’s assumption of a single failing station by adding a masking layer that randomly drops input stations during training (similar to dropout) and a gating mechanism that estimates the reliability of each station’s data in real time. This enables the model to operate when multiple stations fail simultaneously.
-
Capability: The system can now simulate outputs even when 20–50% of the sensor network is offline, by learning robust spatial dependencies from partial observations—critical for disaster response or degraded infrastructure.
- Static Context Injection via Cross-Modal Embeddings
-
Improvement: Add a static-data module that encodes non-temporal features (e.g., building density, road proximity, elevation) using a graph neural network (GNN) or a learned embedding layer. This module outputs a context vector that is concatenated with the decoder’s hidden states, following SATADL’s pattern of combining local and global representations.
-
Capability: The system can incorporate urban planning data, topographical maps, or facility layouts to improve prediction accuracy in regions with sparse historical data, making it transferable to new cities with limited monitoring infrastructure.
- Uncertainty-Aware Forecasting with Quantile Outputs
-
Improvement: Replace the MSE loss with a quantile regression loss (e.g., pinball loss) and output prediction intervals alongside point estimates. Use the temporal attention weights to estimate epistemic uncertainty, and the residual variance from the biLSTM decoder to estimate aleatoric uncertainty.
-
Capability: The system provides confidence bounds for each simulated measurement, enabling risk-aware decision-making (e.g., triggering health alerts only when the upper bound exceeds a safety threshold) and flagging low-confidence predictions for manual inspection.
- Cross-Domain Transfer Learning for Sensor Networks
-
Improvement: Pre-train SATADL’s spatial and temporal attention modules on a large, multi-city dataset (e.g., all four cities combined) with a domain-adversarial training objective to remove city-specific biases. Then fine-tune on a new city with only a few weeks of data.
-
Capability: The system can be deployed in new urban areas within days, achieving near-SOTA performance with minimal local data—useful for rapidly expanding smart-city initiatives or temporary monitoring during events.
- Real-Time Anomaly Detection and Self-Healing
-
Improvement: Add a monitoring layer that compares SATADL’s predictions against actual incoming data (when available) to detect sensor drift or sudden malfunctions. If the prediction error exceeds a threshold, the system automatically re-trains the temporal attention weights on recent data using online learning.
-
Capability: The system not only simulates missing data but also acts as a diagnostic tool, identifying which stations are degrading before they fail, and continuously adapting to slow environmental changes (e.g., seasonal shifts) without full retraining.
- Multi-Pollutant Joint Prediction
-
Improvement: Extend the output layer to predict multiple pollutants simultaneously (e.g., PM10, NO2, O3) by sharing the spatial and temporal attention modules but using separate decoder heads with a covariance-aware loss function (e.g., negative log-likelihood of a multivariate Gaussian).
-
Capability: The system captures cross-pollutant chemical interactions (e.g., O3 formation from NO2), improving accuracy for all pollutants and enabling health impact assessments that require combined exposure metrics.
- Edge-Deployable Lightweight Variant
-
Improvement: Distill SATADL into a smaller student model using knowledge distillation, pruning attention heads, and quantizing weights to 8-bit integers. Replace the biLSTM with a causal transformer decoder to enable parallelization on edge devices.
-
Capability: The system runs on low-power microcontrollers or edge gateways in real time, enabling on-site simulation at each monitoring station without cloud connectivity—ideal for remote or resource-constrained deployments.
What the Improved AI System Can Do:
-
Serve as a universal sensor-network copilot, predicting missing data, detecting faults, and providing uncertainty bounds across air quality, climate, energy, and traffic domains.
-
Operate under partial network failures, integrate static geographic knowledge, and transfer to new cities with minimal retraining.
-
Run in real time on edge hardware, enabling autonomous, self-healing monitoring infrastructure that supports public health alerts, urban planning, and disaster response.
Abstract
Poor air quality in urban areas is driven by a complex chain of processes and presents a significant public health concern. To better understand and control the mechanisms that determine air quality, cities deploy networks of measurement stations, and launch initiatives for collecting denser data about the concentration of pollutants in the atmosphere. Extracting insights from the stations relies on their reliable and uninterrupted operation. However, hardware is susceptible to faults and black- outs that may result in data unavailability, which affects the overall quality of analyses. In this paper, we present a deep-learning model, called SATADL, which can infer complex relations and output multiple-hour-ahead air-quality forecasts. The goal of the model is to simulate the mea- surements of an unresponsive station until its operation is restored. The architecture of the model, which allows it to extract information from different aspects of the data, is described in detail and a careful examination of all of its components is provided. We demonstrate the performance of SATADL on four sets of air quality stations from around the world, by using it to simulate the concentration of PM10 for periods of hypothetical failures of one of the measurement stations, lasting for as long as 48 hours. A selection of baseline and published deep learning models were trained and used as a benchmark. The results show that SATADL per- forms better across different prediction windows, for both coefficient of determination and root mean squared error, demonstrating its suitability as a virtual proxy station.
Sources
- Informer: Beyond Efficient Transformer for Long Sequence Time-Series Forecasting
- Recover Missing Sensor Data with Iterative Imputing Network
- Neural Machine Translation by Jointly Learning to Align and Translate
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks