Benchmarking Time Series Generation Methods for Privacy-Preserving Forecasting

arXiv:2608.10891 · cs.LG · Submitted 2026-08-11 · Read on arXiv

Luis Amorim, Vitor Cerqueira, Moises Santos, Paulo J. Azevedo, Carlos Soares

University of Minho · University of Porto · Laboratory for Artificial Intelligence and Computer Science · Fraunhofer Portugal AICOS · HASLab - INESCTEC · University of Coimbra

cs.LG

Submitted: 2026-08-11

Updated: 2026-08-12

Code: https://github.com/Amorim009/Grasynda

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

The gist: This paper addresses the underexplored question of how well synthetic time series generation methods perform when their output is the sole training source for forecasting models, and how much privacy

Terminology

Summary

This paper addresses the underexplored question of how well synthetic time series generation methods perform when their output is the sole training source for forecasting models, and how much privacy risk the released series carry. The authors note that most synthetic time series generation methods have been developed primarily for data augmentation, where generated series supplement the original training set, but this approach falls short when the original data cannot be used for training due to privacy concerns.

The paper introduces a benchmark evaluating synthetic generation methods and noise-based anonymization baselines under a Train on Synthetic, Test on Real (TSTR) protocol, jointly assessing forecasting performance and distance-based empirical privacy risk across seven datasets. The central research questions are: how do synthetic time series generation methods compare when used as the sole training source in forecasting problems? How private are the released series relative to the original? And does any approach achieve a favorable trade-off between forecasting utility and privacy?

The authors also introduce Grasynda-P, a privacy-motivated extension of the graph-based generator Grasynda, incorporating two modifications: (1) matrix ensembling, which reduces the influence of individual training series on generation by averaging the transition matrix of a target series with its r most similar matrices from the training set using Frobenius distance, and (2) kernel density estimation for value reconstruction, which smooths the mapping from discrete states back to continuous values, preventing the reproduction of exact original values.

The experimental design uses seven publicly available datasets from four forecasting competitions: M1, M3, Tourism, and NN3, spanning industry, demography, and economics domains. The forecasting horizon is 12 for monthly and 8 for quarterly series. The methods evaluated include transformation methods (Scaling, Jitter, M-Warp, T-Warp), pattern mixing methods (DBA, SeasonalMBB, TSMixup), deep generative models (TimeVAE, TSDiff), Grasynda, Grasynda-P, and noise-based baselines (LPA and FPA with ε=1). NHITS is used as the forecasting model, and MASE is the evaluation metric. Privacy is measured using Distance to Closest Record (DCR) and Nearest Neighbour Distance Ratio (NNDR).

The results show that: (1) no generation method fully substitutes for original training data, as training on original data achieves the best average rank; (2) noise-based anonymization yields the strongest privacy but the worst forecasting performance, with LPA and FPA occupying the bottom ranks and producing particularly large MASE values; (3) simple transformation-based generators outperform deep generative models for forecasting in this setting, with methods like Jitter and Scaling outperforming TimeVAE and TSDiff; and (4) Grasynda-P lies on the Pareto frontier, achieving competitive forecasting with stronger privacy separation than other generators.

Regarding the privacy evaluation, noise-based anonymization methods achieve the highest privacy separation. Among synthetic generation methods, Grasynda-P ranks first on DCR, while TSDiff ranks first on NNDR. Grasynda-P achieves higher empirical privacy than Grasynda across all metrics, with statistically significant improvements on DCR across all seven datasets (Wilcoxon signed-rank test, p < 0.05) and NNDR improvements significant on five of seven datasets. The forecasting performance difference between Grasynda-P and Grasynda is not statistically significant.

The trade-off analysis reveals a systematic relationship between forecasting utility and privacy separation, with a Spearman rank correlation of ρ = -0.776 (p = 0.002) between average forecasting rank and average privacy rank across all methods, and ρ = -0.638 (p = 0.035) when restricting to synthetic generators only. Grasynda-P is not dominated by any other synthetic generator on both objectives simultaneously, while Grasynda is Pareto-dominated by Grasynda-P.

The authors discuss several limitations: empirical privacy is measured through distance-based metrics that do not provide formal differential privacy guarantees; the benchmark covers only univariate monthly and quarterly series; a single forecasting architecture (NHITS) is used; and the two modifications defining Grasynda-P are evaluated jointly without isolating individual contributions. The paper concludes that graph-based generation via Grasynda-P achieves a favorable position on the Pareto frontier, suggesting that operating on transition structures rather than direct value transformations offers a promising design principle for privacy-aware time series synthesis.

Improvements for AI systems

Improvements to AI Systems:

  1. Privacy-Aware Synthetic Data Generation for Forecasting Models
  • Implement a graph-based generative module (Grasynda-P) that uses matrix ensembling and kernel density estimation to produce synthetic time series with reduced memorization of original training data.

  • The improved system can generate training datasets that preserve temporal dynamics while minimizing exact value leakage, enabling safe release of synthetic data for collaborative forecasting without exposing raw sensitive records.

  1. Pareto-Optimal Trade-Off Optimization in Data Synthesis
  • Integrate a multi-objective optimization layer that balances forecasting utility (e.g., MASE) and privacy separation (e.g., DCR, NNDR) using the observed Spearman correlation (ρ = -0.776) as a constraint.

  • The improved system can automatically select or tune generation methods (e.g., preferring Grasynda-P over deep generative models like TimeVAE/TSDiff) to achieve a user-specified privacy-utility operating point, rather than relying on ad-hoc choices.

  1. Benchmark-Driven Method Selection for TSTR Protocols
  • Build a recommendation engine that, given a new time series dataset (univariate, monthly/quarterly), predicts the best synthetic generation method based on the benchmark results (e.g., transformation methods like Jitter/Scaling over deep models).

  • The improved system can preemptively avoid low-performing methods (e.g., noise-based LPA/FPA) and guide practitioners toward generators that maintain forecasting accuracy when original data is unavailable, reducing trial-and-error costs.

  1. Privacy Risk Auditing for Synthetic Time Series
  • Embed a post-hoc privacy assessment module that computes DCR and NNDR on generated series relative to the original training set, flagging outputs with low separation (e.g., below Grasynda-P’s thresholds).

  • The improved system can automatically reject or refine synthetic datasets that pose high re-identification risk, providing a practical safeguard before data release.

  1. Hybrid Generation Strategy Combining Transformation and Graph-Based Methods
  • Design an ensemble that switches between simple transformations (e.g., Jitter for high utility) and Grasynda-P (for high privacy) based on the desired trade-off, leveraging the finding that no single method dominates across all objectives.

  • The improved system can dynamically adapt to dataset characteristics (e.g., series length, domain) to maximize forecasting performance while maintaining a user-defined privacy floor, outperforming static choices.

  1. Formal Privacy Guarantee Integration
  • Extend Grasynda-P’s empirical privacy (distance-based) with differential privacy mechanisms (e.g., calibrated noise on transition matrices) to provide formal guarantees, addressing the paper’s noted limitation.

  • The improved system can offer provable privacy bounds (e.g., ε-DP) while retaining the observed forecasting utility, making it suitable for regulated domains like healthcare or finance where empirical metrics alone are insufficient.

Abstract

Time series forecasting in privacy-sensitive domains often requires training models on released data rather than original observations. Synthetic time series generation has been developed primarily for data augmentation, where generated series supplement the original training set. How well these methods perform when fully replacing the original data - and how much privacy risk the released series carry - remains underexplored. We address this gap through a benchmark evaluating synthetic generation methods and noise-based anonymization baselines under a Train on Synthetic, Test on Real (TSTR) protocol. We jointly assess forecasting performance and distance-based empirical privacy risk across seven datasets, characterizing the trade-off between these objectives. We also introduce Grasynda-P, a privacy-motivated extension of the graph-based generator Grasynda, incorporating matrix ensembling and kernel density estimation. Our results show that: (1) no generation method fully substitutes for original training data; (2) noise-based anonymization yields the strongest privacy but the worst forecasting performance; (3) simple transformation-based generators outperform deep generative models for forecasting in this setting; and (4) Grasynda-P lies on the Pareto frontier, achieving competitive forecasting with stronger privacy separation than other generators. This benchmark establishes a reference point for evaluating and developing new privacy-aware synthetic time series generation methods.

Related papers