A Critical Audit of Spatiotemporal Forecasting Benchmark Datasets and Models
summary
The gist
Graph neural networks (GNNs) are routinely employed for short-range forecasting on multivariate time series with a spatial graph structure, but this work critically audits widely adopted benchmark
In short
This research audits popular forecasting benchmarks to find biases. It analyzes correlations and critiques evaluation methods, revealing that many studies use flawed protocols like differencing. The findings suggest spatially-unaware linear models often outperform complex Graph Neural Networks on certain datasets, demanding more rigorous statistical comparison.
Key concepts
- Statistical Analysis of Benchmark Datasets
- The study measures temporal and spatial correlations across five datasets. It found that some datasets, like Chickenpox, show almost no spatial information at various lags, while traffic data exhibits strong, long-lasting temporal and spatial memory.
- Critique of Evaluation Protocols
- Many studies incorrectly compare models by differencing raw time series. Undifferencing reveals that the temporal correlation in datasets like Chickenpox is persistent for up to ten lags and gains significant linear informativeness, suggesting undifferenced data allows better generalization.
- SARIMA Residual Training
- The authors propose training Graph Neural Networks (GNNs) on the residuals of SARIMA models. This approach successfully raises GNN performance to state-of-the-art levels on traffic datasets, showing a promising combination of traditional time series and spatial learning.
Terminology used across episodes
This episode discusses
- A Critical Audit of Spatiotemporal Forecasting Benchmark Datasets and Models · Paper Radio
- Message-Passing State-Space Models: Improving Graph Learning with Modern Sequence Modeling
- Graph Neural Ordinary Differential Equations
- Machine Learning in Python: Main developments and technology trends in data science, machine learning, and artificial intelligence
- Beyond Homophily in Graph Neural Networks: Current Limitations and Effective Designs
The paper
A Critical Audit of Spatiotemporal Forecasting Benchmark Datasets and Models · Read on arXiv
Imperial College London · Ruhr University Bochum
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "A Critical Audit of Spatiotemporal Forecasting Benchmark Datasets and Models".
Jane: Graph neural networks (GNNs) are routinely employed for short-range forecasting on multivariate time series with a spatial graph structure,
Tom: First, who's behind it and why it matters.
Paper summary: Tom: So, focusing on that summary we just heard, Tom says this paper critically audits five major benchmark datasets: Chickenpox, PedalMe, WikiMaths, METR-LA, and PEMS-BAY. The core thesis is that analyzing these benchmarks using classical time series methods reveals why spatially-unaware linear models actually perform quite well against what was previously thought to be superior GNNs.
Jane: It also points out that the evaluation protocols themselves have some shortcomings; for instance, many studies use differenced time series instead of raw data, which the authors find leads to inferior predictive performance across the board. They highlight that undoing this differencing on Chickenpox reveals a much more persistent temporal correlation and increased spatial informativeness.
Lu: It’s interesting how they tie the temporal signal preservation directly to the spatial signal; it seems like simply cleaning up the data preprocessing can fundamentally change how useful that information is for prediction.
Meng: From an engineering standpoint, if we are training models on differenced data when raw data shows better generalization metrics, that suggests our current testing pipeline might be training us on suboptimal inputs.
Lalam: This audit matters because it encourages a more rigorous statistical evaluation process rather than just comparing model complexity against these specific benchmarks. It pushes us to think about the underlying signal structure first.
Conclusion: Tom: Wrapping up this discussion on "A Critical Audit of Spatiotemporal Forecasting Benchmark Datasets and Models," the authors are essentially saying we need to be more careful about how we benchmark our forecasting models. They suggest that benchmarking against these specific datasets might not give us the full picture because the temporal signals in them drive a lot of the observed performance differences.
Jane: I agree, Tom, and they specifically recommend keeping Chickenpox and PedalMe undifferenced when used for benchmarking purposes because that helps preserve those low-frequency signals. They also found that simple linear models like AR(H) per node and DLinear can actually compete very well on some of these datasets.
Lu: The implication here is significant: we might be overvaluing the spatial complexity captured by GNNs when, in reality, the temporal structure alone is doing a lot of the heavy lifting for these specific benchmark problems.
Meng: If we adopt their suggestion to use SARIMA residuals as a training target for GNNs on traffic data like METR-LA and PEMS-BAY, that seems like a practical way to combine the strengths of both approaches effectively.
Lalam: It suggests that instead of just chasing the most complex architecture, we should focus more on understanding the temporal and spatial correlations statistically before jumping into deep learning structures for every single problem.
Tom: So, to summarize this audit: don't be afraid to question your benchmarks and consider how simple models might actually hold up against the hype. That’s a big message from this work by Kenneth Martin et al.
Jane: Exactly, Tom; it’s about moving toward more rigorous statistical evaluation so we can build better forecasting systems for real-world applications.
More episodes
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language
- 2508.08833-An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems
- 2405.04118-Policy Learning with a Language Bottleneck