msData: A Millisecond-Resolution Network Dataset for Advancing Time Series Foundation Models

arXiv:2603.16497 · cs.LG, cs.AI · Submitted 2026-03-17 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "msData: A Millisecond-Resolution Network Dataset for Advancing Time Series Foundation Models".

Jane: The paper was written by Subina Khanal, Seshu Tirupathi, Merim Dzaferagic, Marco Ruffini and Torben Bach Pedersen from Department of Computer Science, Aalborg University, Aalborg, Denmark and IBM Research Europe, Dublin, Ireland and Trinity College Dublin, The University of Dublin, Dublin, Ireland..

Tom: Stay tuned as we take you through the paper and discuss its implications.

The Core Contribution: Tom: So, we have msData, but what exactly is this new dataset? The authors are detailing a real-world capture of 5G Radio Access Network performance measurements.

Jane: It’s essentially taking the operational data from an Open Ireland testbed where we can actually see how the network behaves in a highly dynamic environment. Instead of using aggregated averages, we're looking at specific physical and medium access control layer features.

Lu: The detail on the Open Radio Access Network is important because it shows that this isn't just a theoretical exercise; it’ has real-world, modular components like the Central Unit and Distributed Unit working together.

Meng: And what makes this dataset unique compared to other datasets like ETTh1 or Traffic—it’s the combination of high frequency and a new domain entirely, right?

Lalam: It shifts our focus from traditional domains like energy or finance to the complex, unpredictable world of wireless connectivity. That is a big cultural shift in how we view data sources for AI training.

Tom: And Jane mentioned that this dataset captures diverse mobility profiles—static, pedestrian, car, bus and train—which really adds depth to the complexity.

Jane: Yes, and it includes traffic from both benign applications like VoIP and even malicious activities such as DDoS-Ripper flows. This makes it a comprehensive picture of network usage.

Lu: We are seeing a model that is built not only on predictable cycles but also on real chaos, which is crucial for robustness in AI systems.

Meng: Practically, this means the engineers building future networks will have data that actually reflects how people move through and use those services moment-to-moment.

Lalam: The inclusion of the unique identifier for each user helps track individual experience, moving beyond a collective average to something far more granular about how individuals interact with infrastructure.

Applications and Challenges: Tom: msData isn't just about collecting data; the authors are suggesting very specific use cases, especially for short-term forecasting.

Jane: They are looking at prediction horizons that start as short as one millisecond up to ninety-six milliseconds, which is incredibly fine-grained.

Lu: This suggests that AI systems built on this data will be able to predict network states with almost instantaneous foresight.

Meng: I am concerned about the volatility; since we are talking about bitrates and channel quality, how do these models handle sudden spikes or rapid fluctuations in load?

Lalam: The structure of the data itself—it has irregular bursts and heavy-tailed dynamics—is designed to train AI to anticipate rapid shifts rather than just smooth trends.

Tom: And when we look at the applications, it's not just predicting throughput; we' are talking about using that prediction to proactively steer users or support load-aware handovers.

Jane: Exactly, Tom; predicting channel quality allows the network to make intelligent decisions before a user even notices their streaming video stutter.

Lu: It also supports identifying malicious activity early on, which is vital for security monitoring systems that rely on behavioral patterns rather than packet inspection.

Meng: From an engineering standpoint, this helps us build more reliable Quality of Service policies because we can anticipate when things are about to change.

Lalam: The ability to predict traffic class transitions means the AI can start allocating resources preemptively, ensuring a smoother user experience across diverse scenarios.

The Benchmarking Results: Tom: This brings us to the benchmarking phase where the authors test traditional machine learning models against TSFMs using this high-frequency data.

Jane: The findings are quite stark; they show that most of the current TSFM configurations perform poorly on this new data distribution, whether you're using them in a zero-shot or fine-tuned setting.

Lu: This confirms our hypothesis that because these models were trained on slow, low-frequency data, they simply lack the necessary internal mechanisms to adapt to the rapid shifts in msData.

Meng: It’s fascinating that even with fine-tuning—which usually helps AI generalize—the performance is suboptimal; it suggests the problem isn't just poor training but a fundamental mismatch in scale.

Lalam: This failure highlights a crucial limitation of current AI paradigms, forcing us to rethink how we structure our pre-training phase to match real-world complexity.

Tom: And while TSFMs are struggling, the authors found that Adaptive Random Forest, or ARF, consistently outperformed all other shallow models and TSFMs in both univariate and multivariate settings.

Jane: It's a powerful validation of a more dynamic approach; ARF is built to handle concept drift—those sudden data shifts—which is exactly what our 5G network data looks like.

Lu: The ability ARF has to dynamically update its ensemble makes it much better suited for handling volatility than static methods, which are just not designed for this kind of erratic behavior.

Meng: If we need reliability in a mission-critical system, the fact that ARF handles the irregular spikes so well is a massive practical advantage over the TSFMs.

Lalam: This suggests that moving beyond purely transformer architectures might be necessary to fully capture the chaotic but predictable nature of high-speed real-world data.

Conclusion and Impact: Tom: So, we've seen how msData bridges a massive data gap, tested the performance of various models, and identified a clear need to advance TSFMs.

Jane: This work is a powerful call for incorporating high-resolution datasets into the training phase to boost the accuracy and generalizability of all time series AI.

Lu: It's not just an improvement; it's an evolution, forcing us to rethink the very definition of what 'generalized AI' looks like in high-speed environments.

Meng: From a deployment perspective, this allows for much tighter control and more robust decision-making when we are dealing with milliseconds of data.

Lalam: This ultimately leads to a more sophisticated cultural shift where our AI systems aren't just statistical tools but genuine predictors of dynamic reality, enabling better societal control over infrastructure.

Tom: We’ve covered so many ground today, from the specific features in this dataset to the future research directions.

Jane: It really is a breakthrough moment for time series modeling, giving us a clear path forward for msData.

Lu: The path is open for true adaptability, allowing us to train models that can handle everything from steady traffic flows to sudden bursts of network congestion.

Meng: We need these tools that can actually adapt to the messy reality of day-to-day operations; we can't afford more static solutions.

Lalam: I hope this research allows our future AI systems to better reflect the dynamic complexity of our global networks, ultimately enhancing human interaction with technology through reliable and fast service.

Tom: Thank you all for helping us unpack "msData: A Millisecond-Resolution Network Dataset for Advancing Time Series Foundation Models."

Jane: It’s been a fantastic conversation; keep an eye on this work as it truly sets the stage for what’s next in time series AI.

cs.LG, cs.AI

Submitted: 2026-03-17

Updated: 2026-08-25

Importance score: 79/100

The gist: The paper introduces "msData," a millisecond-resolution network dataset specifically designed to address limitations in current Time Series Foundation Models (TSFMs).

Key concepts

msData
A new dataset capturing real-world 5G Radio Access Network performance. It includes operational data from an Open Ireland testbed, focusing on physical and medium access control layer features. It captures diverse mobility profiles (static, pedestrian, car) and both benign and malicious network activity.
Time Series Foundation Models (TSFMs)
AI models designed to learn patterns from sequential data. The discussion highlights that current TSFM configurations perform poorly on high-frequency datasets like msData because they lack the internal mechanisms to adapt to rapid shifts in network conditions.
Adaptive Random Forest (ARF)
A machine learning model identified as a strong performer against TSFMs. ARF is specifically designed to handle concept drift or sudden data shifts, making it highly effective for volatile 5G network data where traditional static methods fail.

Terminology

Summary

The paper introduces msData, a millisecond-resolution network dataset specifically designed to address limitations in current Time Series Foundation Models (TSFMs). By providing high-granularity data, msData enables advanced research into complex time series prediction tasks, particularly within the domain of network traffic analysis. The comprehensive evaluation presented demonstrates that this dataset is highly suitable for training and benchmarking various state-of-the-art models across both univariate and multivariate settings.

Dataset Feature Space and Scope

The network data contains 47 features in total, providing enough features for multivariate setting. This feature richness ensures that the dataset is well-suited for training TSFMs, similar to existing pre-trained datasets. In initial experiments, the researchers utilized a subset of four important features from the network dataset. Furthermore, the analysis was extended to include ten features in total, allowing for an evaluation of benchmarked models in a fully multivariate setting.

Evaluation of Generalization and Transfer Learning

The performance of the benchmarked models is rigorously tested on filtered data that represents mobility patterns and traffic classes that differ from those presented in Section 3.2. This setup is critical for assessing the dataset's potential for transfer learning; specifically, by training models on one set of mobility patterns and traffic classes and evaluating them on a different set, we can assess how well knowledge learned in one context generalizes to another.

Performance comparisons reveal distinct trends:

  • In the univariate setting, TTM was observed to outperform ARF.

  • However, when evaluated in the multivariate setting, ARF achieves better performance compared to the other models.

  • Further evaluation across additional filtered subsets of data consistently showed that TTM had poorer performance compared to ARF in most traffic labels, indicating the stronger generalization of ARF in the multivariate setting.

Hyper-Parameter Tuning and Model Benchmarking

The study details extensive hyper-parameter tuning for both TTM and Lag-Llama models. For TTM, initial testing revealed that a learning rate of 0.00001 achieved the lowest errors. For the main results, an optimal learning rate of 0.0011 was used, resulting in an RMSE of 0.0391 and an MAE of 0.0249. The researchers noted that increasing the number of epochs did not significantly change performance beyond default settings.

For Lag-Llama, initial evaluation found that a context length of 25 performed better than context length 15, although a context length of 5 was used for the main results to ensure a fair comparison with the shallow models. The optimal batch size was determined to be 32. Across all model comparisons, ARF consistently maintained superior performance:

  • ARF continues to outperform TTM in the multivariate setting.

  • Despite Chronos-base having a more than fourfold increase in model size compared to Chronos-small, the improvement over Chronos-small appeared modest, and ARF continues to outperform Chronos in terms of RMSE, suggesting the performance gap is not solely due to model scale.

Improvements for AI systems

(Note to Self: Extreme diligence required. Do not generalize findings; propose architectural or methodological enhancements based on observed limitations in the current study.)

Based on a meticulous review of this paper detailing high-frequency network bitrate prediction using Time Series Forecasting Models (TSFMs), I have identified several critical areas for improvement. The current work establishes strong benchmarks, particularly highlighting the superior performance of ARF in multivariate settings and providing valuable hyperparameter tuning insights. However, significant opportunities exist to enhance robustness, generalization, and real-time applicability of these systems.

Here are the specific improvements I propose for the AI system:


  • Problem Addressed: The current study notes using a subset of four important features initially, and later expanding to ten. While multivariate analysis is performed, the models (especially TTM/Lag-Llama) treat all input features equally after selection. The importance of features can change drastically based on the network state (e.g., congestion vs. idle).

  • Improvement: Implement a Self-Attention Mechanism or Graph Neural Network (GNN) layer immediately following the feature embedding stage, before feeding data into the core TSFM architecture (like Lag-Llama). This mechanism must dynamically assign weights to the input features at each timestep based on their observed correlation with high prediction error in recent history.

  • Technical Specification: The system should incorporate a weighted average attention score alpha t for each feature f i at time t, such that the input embedding E't = sum i=1 N alpha t, i E(f i). This makes the model feature-aware rather than just multi-variate.

  • Problem Addressed: Network traffic patterns are non-stationary; network behavior changes over time (concept drift). The current evaluation is static, assuming the learned knowledge remains valid.

  • Improvement: Integrate a real-time monitoring module using statistical process control (e.g., CUSUM or EWMA charts) that continuously monitors the prediction error distribution and the input feature distributions against the training baseline. If the model's prediction error exceeds a statistically determined threshold (sigma drift) for a sustained period, or if key feature distributions shift significantly, a retraining trigger must be activated.

  • Technical Specification: The system should transition from standard inference mode to Adaptive Learning Mode, initiating either incremental fine-tuning on the most recent window of data or flagging the need for full model retraining with human supervision.

  • Problem Addressed: The paper demonstrates potential for transfer learning (training on one mobility pattern/traffic class and evaluating on another). However, the current comparison is merely an evaluation of performance; it lacks a structured mechanism to quantify what knowledge is transferable.

  • Improvement: Modify the TSFM architecture (e.g., Lag-Llama) to explicitly separate Domain-Specific Layers (handling traffic class/mobility pattern details) from General Domain Layers (handling universal temporal dependencies, like diurnal cycles or general congestion effects). During transfer, the system should first freeze the weights of the General Domain Layers and only fine-tune the Domain-Specific Layers using a minimal amount of target data.

  • Technical Specification: This requires implementing an Adversarial Domain Adaptation loss function. The model must be trained not only to minimize prediction error but also to maximize its inability for a domain discriminator network to tell which domain (source or target) the input sample came from, thus forcing feature representations to become domain-agnostic and maximally transferable.

  • Problem Addressed: The paper compares multiple models (ARF, TTM, Lag-Llama). While ARF is noted as superior, relying on a single best model is risky in mission-critical systems. Furthermore, the reported metrics (RMSE/MAE) are point estimates and do not provide confidence intervals.

  • Improvement: Develop an Ensemble Forecasting System that combines the outputs of the top two performing models (e.g., ARF and Lag-Llama). Crucially, this ensemble must be coupled with a Bayesian Neural Network (BNN) approach to generate predictive uncertainty (sigma pred) alongside the mean prediction mu pred.

  • Technical Specification: The final output should be a tuple (mu pred, sigma pred). This allows downstream systems (e.g., network resource allocators) to make risk-aware decisions: if sigma pred is high, the system knows its prediction is unreliable and must trigger a fallback mechanism or alert an operator, rather than blindly trusting the average prediction.

By implementing these improvements, the resulting AI system transforms from a static predictive tool into a Cognitive Network Resource Management Engine.

  1. Proactive Congestion Mitigation: The system can predict not just what the bitrate will be, but how certain it is about that prediction. If high predicted bitrates are coupled with high uncertainty (sigma pred), it can proactively suggest load shedding or rerouting paths to prevent potential bottlenecks before they occur, minimizing service disruption costs.

  2. Zero-Shot Network Adaptation: The system gains the ability to maintain peak performance when deployed into entirely new network segments or traffic classes (e.g., switching from analyzing enterprise VPN traffic to analyzing IoT sensor data). It achieves this by automatically identifying and leveraging universal temporal patterns (General Domain Layers) while minimizing the required retraining data for localized context adaptation.

  3. Resource Allocation Optimization: Network engineers can use the system's output (mu pred, sigma pred) to optimize resource provisioning. Instead of allocating resources based only on mean predicted demand, they allocate based on Value at Risk (VaR), ensuring that the probability of failure due to under-provisioning remains below a critical threshold defined by sigma pred.

  4. Automated Performance Auditing: The integrated drift detection module allows the system to function as its own quality assurance layer. It continuously audits its own performance, automatically alerting human operators when model degradation is detected, thereby drastically reducing the Mean Time To Recovery (MTTR) following unforeseen network changes or cyber-attacks.

Sources

Related papers