msData: A Millisecond-Resolution Network Dataset for Advancing Time Series Foundation Models
summary
The gist
The paper introduces "msData," a millisecond-resolution network dataset specifically designed to address limitations in current Time Series Foundation Models (TSFMs).
In short
The episode discusses 'msData,' a millisecond-resolution network dataset capturing real-world 5G Radio Access Network performance. Hosts analyze how this high-frequency data, compared to traditional datasets, enables AI models to predict network states and handle dynamic conditions. They conclude that current Time Series Foundation Models (TSFMs) struggle with this rapid change, suggesting a need for more adaptive approaches like Adaptive Random Forest.
Key concepts
- msData
- A new dataset capturing real-world 5G Radio Access Network performance. It includes operational data from an Open Ireland testbed, focusing on physical and medium access control layer features. It captures diverse mobility profiles (static, pedestrian, car) and both benign and malicious network activity.
- Time Series Foundation Models (TSFMs)
- AI models designed to learn patterns from sequential data. The discussion highlights that current TSFM configurations perform poorly on high-frequency datasets like msData because they lack the internal mechanisms to adapt to rapid shifts in network conditions.
- Adaptive Random Forest (ARF)
- A machine learning model identified as a strong performer against TSFMs. ARF is specifically designed to handle concept drift or sudden data shifts, making it highly effective for volatile 5G network data where traditional static methods fail.
Terminology used across episodes
This episode discusses
- msData: A Millisecond-Resolution Network Dataset for Advancing Time Series Foundation Models · Paper Radio
- GIFT-Eval: A Benchmark For General Time Series Forecasting Model Evaluation
- Diffusion Models for Time Series Forecasting: A Survey
The paper
msData: A Millisecond-Resolution Network Dataset for Advancing Time Series Foundation Models · Read on arXiv
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "msData: A Millisecond-Resolution Network Dataset for Advancing Time Series Foundation Models".
Jane: The paper was written by Subina Khanal, Seshu Tirupathi, Merim Dzaferagic, Marco Ruffini and Torben Bach Pedersen from Department of Computer Science, Aalborg University, Aalborg, Denmark and IBM Research Europe, Dublin, Ireland and Trinity College Dublin, The University of Dublin, Dublin, Ireland..
Tom: Stay tuned as we take you through the paper and discuss its implications.
The Core Contribution: Tom: So, we have msData, but what exactly is this new dataset? The authors are detailing a real-world capture of 5G Radio Access Network performance measurements.
Jane: It’s essentially taking the operational data from an Open Ireland testbed where we can actually see how the network behaves in a highly dynamic environment. Instead of using aggregated averages, we're looking at specific physical and medium access control layer features.
Lu: The detail on the Open Radio Access Network is important because it shows that this isn't just a theoretical exercise; it’ has real-world, modular components like the Central Unit and Distributed Unit working together.
Meng: And what makes this dataset unique compared to other datasets like ETTh1 or Traffic—it’s the combination of high frequency and a new domain entirely, right?
Lalam: It shifts our focus from traditional domains like energy or finance to the complex, unpredictable world of wireless connectivity. That is a big cultural shift in how we view data sources for AI training.
Tom: And Jane mentioned that this dataset captures diverse mobility profiles—static, pedestrian, car, bus and train—which really adds depth to the complexity.
Jane: Yes, and it includes traffic from both benign applications like VoIP and even malicious activities such as DDoS-Ripper flows. This makes it a comprehensive picture of network usage.
Lu: We are seeing a model that is built not only on predictable cycles but also on real chaos, which is crucial for robustness in AI systems.
Meng: Practically, this means the engineers building future networks will have data that actually reflects how people move through and use those services moment-to-moment.
Lalam: The inclusion of the unique identifier for each user helps track individual experience, moving beyond a collective average to something far more granular about how individuals interact with infrastructure.
Applications and Challenges: Tom: msData isn't just about collecting data; the authors are suggesting very specific use cases, especially for short-term forecasting.
Jane: They are looking at prediction horizons that start as short as one millisecond up to ninety-six milliseconds, which is incredibly fine-grained.
Lu: This suggests that AI systems built on this data will be able to predict network states with almost instantaneous foresight.
Meng: I am concerned about the volatility; since we are talking about bitrates and channel quality, how do these models handle sudden spikes or rapid fluctuations in load?
Lalam: The structure of the data itself—it has irregular bursts and heavy-tailed dynamics—is designed to train AI to anticipate rapid shifts rather than just smooth trends.
Tom: And when we look at the applications, it's not just predicting throughput; we' are talking about using that prediction to proactively steer users or support load-aware handovers.
Jane: Exactly, Tom; predicting channel quality allows the network to make intelligent decisions before a user even notices their streaming video stutter.
Lu: It also supports identifying malicious activity early on, which is vital for security monitoring systems that rely on behavioral patterns rather than packet inspection.
Meng: From an engineering standpoint, this helps us build more reliable Quality of Service policies because we can anticipate when things are about to change.
Lalam: The ability to predict traffic class transitions means the AI can start allocating resources preemptively, ensuring a smoother user experience across diverse scenarios.
The Benchmarking Results: Tom: This brings us to the benchmarking phase where the authors test traditional machine learning models against TSFMs using this high-frequency data.
Jane: The findings are quite stark; they show that most of the current TSFM configurations perform poorly on this new data distribution, whether you're using them in a zero-shot or fine-tuned setting.
Lu: This confirms our hypothesis that because these models were trained on slow, low-frequency data, they simply lack the necessary internal mechanisms to adapt to the rapid shifts in msData.
Meng: It’s fascinating that even with fine-tuning—which usually helps AI generalize—the performance is suboptimal; it suggests the problem isn't just poor training but a fundamental mismatch in scale.
Lalam: This failure highlights a crucial limitation of current AI paradigms, forcing us to rethink how we structure our pre-training phase to match real-world complexity.
Tom: And while TSFMs are struggling, the authors found that Adaptive Random Forest, or ARF, consistently outperformed all other shallow models and TSFMs in both univariate and multivariate settings.
Jane: It's a powerful validation of a more dynamic approach; ARF is built to handle concept drift—those sudden data shifts—which is exactly what our 5G network data looks like.
Lu: The ability ARF has to dynamically update its ensemble makes it much better suited for handling volatility than static methods, which are just not designed for this kind of erratic behavior.
Meng: If we need reliability in a mission-critical system, the fact that ARF handles the irregular spikes so well is a massive practical advantage over the TSFMs.
Lalam: This suggests that moving beyond purely transformer architectures might be necessary to fully capture the chaotic but predictable nature of high-speed real-world data.
Conclusion and Impact: Tom: So, we've seen how msData bridges a massive data gap, tested the performance of various models, and identified a clear need to advance TSFMs.
Jane: This work is a powerful call for incorporating high-resolution datasets into the training phase to boost the accuracy and generalizability of all time series AI.
Lu: It's not just an improvement; it's an evolution, forcing us to rethink the very definition of what 'generalized AI' looks like in high-speed environments.
Meng: From a deployment perspective, this allows for much tighter control and more robust decision-making when we are dealing with milliseconds of data.
Lalam: This ultimately leads to a more sophisticated cultural shift where our AI systems aren't just statistical tools but genuine predictors of dynamic reality, enabling better societal control over infrastructure.
Tom: We’ve covered so many ground today, from the specific features in this dataset to the future research directions.
Jane: It really is a breakthrough moment for time series modeling, giving us a clear path forward for msData.
Lu: The path is open for true adaptability, allowing us to train models that can handle everything from steady traffic flows to sudden bursts of network congestion.
Meng: We need these tools that can actually adapt to the messy reality of day-to-day operations; we can't afford more static solutions.
Lalam: I hope this research allows our future AI systems to better reflect the dynamic complexity of our global networks, ultimately enhancing human interaction with technology through reliable and fast service.
Tom: Thank you all for helping us unpack "msData: A Millisecond-Resolution Network Dataset for Advancing Time Series Foundation Models."
Jane: It’s been a fantastic conversation; keep an eye on this work as it truly sets the stage for what’s next in time series AI.
More episodes
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language