Towards Generalisable Time Series Understanding Across Domains
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Towards Generalisable Time Series Understanding Across Domains".
Jane: The paper was written by Özgün Turgut, Philip Müller, Martin J Menten and Daniel Rueckert from Technical University of Munich and Klinikum rechts der Isar, Technical University of Munich and Munich Center for Machine Learning and Imperial College London.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back, everyone! Tom here with my co-host Jane, and we are digging into a fresh arXiv paper that's got a mouthful of a title — “Towards Generalisable Time Series Understanding Across Domains.” Jane, what do you make of that title before we even open the hood?
Jane: Tom, I love it because it’s honest. It says “towards,” which tells you this is a stepping stone, not the final answer. And “across domains” is the whole ballgame — time series data is everywhere, from heart monitors to weather stations to stock tickers, and they all look different. The authors are basically saying, we want one model that can handle all of them.
Tom: Exactly. And that’s a big deal because most time series models are trained on one domain and they fail when you switch. You train on EEG brainwaves, and then you try to use it on audio, and it just crumbles.
Jane: Right. So the title is promising something ambitious — a model that understands time series in general, not just one flavor. And the authors are from TU Munich and Imperial College London, so you know there’s serious engineering behind this.
Tom: Yeah, and they’ve named the model OTiS — Open model for general Time series analysis. Catchy, right? But the real question is, how do you even start building something like that when the data is so messy and different everywhere?
Jane: That’s the million-dollar question, and I think the title hints at their answer — they don’t force everything into one box. They let each domain keep its own identity, and that’s the clever part we’re about to unpack.
Tom: And I can’t wait to get into that, because if they pulled this off, it changes the game for a lot of industries. Stick around — we’re just getting started with “Towards Generalisable Time Series Understanding Across Domains.”
Summary: Tom: Alright Jane, we’re back with “Towards Generalisable Time Series Understanding Across Domains.” Give us the quick version — what did these folks actually do?
Jane: So Tom, they built a model that can be pre-trained on a giant pile of time series data from eight different domains — medicine, engineering, natural sciences, finance — and then fine-tuned for specific tasks like classification, regression, and forecasting. The whole corpus has over six hundred forty thousand samples and eleven billion time points.
Tom: eleven billion — that’s not a typo, folks. And they’re not just throwing all that data into one blender. They designed a tokenizer that respects the differences between domains. Like, an ECG has twelve leads that relate to each other in a specific physical way, while a weather station has twenty-one sensors that interact differently.
Jane: Exactly. And that’s the key insight — previous models tried to standardize everything, like resampling all data to the same frequency or chopping it into uniform chunks. That destroys the relationships between variables and the temporal patterns that make each domain unique. OTiS instead learns a “domain signature” — a set of embeddings that capture what makes each domain special.
Tom: And they also added a dual masking strategy during pre-training. Sometimes they randomly mask parts of the input, and sometimes they mask the second half of the time series to force the model to predict the future from the past. That way, the model learns both bidirectional relationships and temporal causality.
Jane: Right, and they didn’t stop there. They added a normalised cross-correlation loss on top of the standard mean squared error. That helps the model capture global patterns, not just local ones. You know, long-range dependencies that matter in things like weather forecasting or financial trends.
Tom: And the results? They benchmarked on fifteen tasks and beat state-of-the-art on ten of them. That’s not a fluke. And they even showed the model has an intrinsic understanding of sine waves — frequency, amplitude, phase — without any fine-tuning. That’s wild.
Jane: It really is. The model seems to have learned something fundamental about time series, not just memorized patterns. And that’s what makes this paper exciting — it’s not just another model, it’s a foundation for a whole new way of thinking about time series analysis.
Tom: So the summary is: they built a model that respects domain differences, learns general time series features, and outperforms specialists. But how did they actually pull that off under the hood? That’s what we’re diving into next.
Improvements: Tom: Back with “Towards Generalisable Time Series Understanding Across Domains.” Jane, we’ve covered the big picture. Now let’s get into the nitty-gritty — what are the actual improvements this paper brings to the table?
Jane: The biggest one, Tom, is the domain-specific tokenizer. Instead of treating every time series the same, they introduce learnable embeddings for each variate in each domain. So for EEG, they have embeddings for each electrode position. For ECG, each lead gets its own embedding. And these embeddings are trained end-to-end with the model.
Tom: And the proof is in the pudding — they did a principal component analysis on these embeddings, and guess what? The EEG embeddings actually reconstruct the three dee spatial layout of the electrodes on the scalp. The ECG embeddings recover Einthoven’s triangle. The model isn’t just learning abstract numbers; it’s learning the physical structure of the data.
Jane: That’s the kind of result that makes you sit up. And the second improvement is the dual masking strategy I mentioned earlier. Random masking teaches the model to fill in gaps anywhere, which builds a rich understanding of inter-variate relationships. Post-fix masking, where you hide the future, teaches the model temporal causality — which is exactly what you need for forecasting.
Tom: And they didn’t just guess at the ratio — they ablated it. seventy-five percent random, twenty-five percent post-fix was the sweet spot. And when they removed either component, performance dropped across the board. That’s solid engineering.
Jane: Then there’s the normalised cross-correlation loss. This is subtle but powerful. The standard MSE loss focuses on getting each point right, but it doesn’t capture whether the overall shape of the prediction matches the ground truth. NCC loss penalizes the model when the trend or the pattern is off, even if individual points are close. That’s how they get those beautiful forecasts you see in the paper.
Tom: And they showed that in the ablation study — removing the NCC loss hurt performance on all tasks, especially forecasting. So it’s not just a nice-to-have; it’s doing real work.
Jane: And one more thing — the scaling study. They showed that more pre-training data consistently improves downstream performance. Model size matters too, but only if you have enough data to feed it. That’s a classic scaling law story, but it’s good to see it confirmed for time series.
Tom: So we’ve got a tokenizer that respects domain structure, a masking strategy that teaches both relationships and causality, and a loss function that captures global patterns. That’s a triple threat. But what does this mean for the real world? That’s where it gets really interesting.
Conclusion: Tom: We’re wrapping up our discussion on “Towards Generalisable Time Series Understanding Across Domains.” Jane, give us the final takeaway.
Jane: The takeaway, Tom, is that this paper shows you can build a genuinely general time series model if you stop pretending all data is the same. By respecting domain differences, teaching the model both relationships and causality, and using a loss that captures global patterns, OTiS achieves state-of-the-art results across classification, regression, and forecasting.
Tom: And the implications are huge. In medicine, where data is scarce and privacy is paramount, you can pre-train on diverse public data and then fine-tune on a small local dataset. The paper even shows that just fine-tuning the domain-specific embeddings — less than 0 point 2k parameters — can produce solid forecasts on unseen frequencies. That’s a game-changer for hospitals with limited data.
Jane: Absolutely. And it’s not just medicine. Finance, energy, climate science — any field that relies on time series data could benefit. The model’s zero-shot capabilities are particularly impressive. It can distinguish between classes it’s never seen before, just from the structure it learned during pre-training.
Tom: And the authors were kind enough to release the code and pre-trained weights, so the community can build on this. That’s how progress happens.
Jane: Right. So we’re saying goodbye to “Towards Generalisable Time Series Understanding Across Domains” — a paper that gives us a foundation for truly general time series analysis. We’ll be watching to see where this goes next.
Tom: And we’re already looking at the next paper on our list. Stay tuned, folks — the conversation continues.
Özgün Turgut, Philip Müller, Martin J Menten, Daniel Rueckert
Technical University of Munich · Klinikum rechts der Isar, Technical University of Munich · Munich Center for Machine Learning · Imperial College London
cs.LG, cs.AI, cs.CV
Submitted: 2026-08-14
Updated: 2026-08-17
Code: https://github.com/oetu/otis
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 71/100
Key concepts
- Domain Signature
- Instead of forcing all time series data into a single format, the model learns unique 'domain signatures.' It uses specialized embeddings that capture the physical structure of each domain—for example, recognizing specific electrode positions in EEG or leads in ECG—allowing it to respect inherent differences between variables.
- Dual Masking Strategy
- This technique involves two methods during pre-training. Random masking helps the model understand relationships between different data points, while post-fix masking forces the model to predict future values based only on past data. This ensures the model learns temporal causality.
- Normalised Cross-Correlation (NCC) Loss
- This loss function goes beyond just matching individual data points. It penalizes the model if the overall trend or pattern of its prediction is incorrect. This allows OTiS to capture long-range, global patterns, which is essential for accurate forecasting.
Terminology
Summary
Summary
This paper introduces a novel pre-training paradigm and model, the open model for general time series analysis (OTiS), designed to address the challenge of heterogeneity in large time series corpora. The authors argue that while foundation models have excelled in natural language processing and computer vision, this potential remains unrealised in time series analysis because existing methods fail to handle the substantial differences in time series characteristics across domains, such as variate count, inter-variate relationships, temporal patterns, and sampling frequency.
To overcome this, the paper proposes three key contributions: a domain-specific tokeniser, a dual masking strategy, and a normalised cross-correlation (NCC) loss. The tokeniser uses learnable domain-specific variate embeddings to account for unique domain characteristics, including inter-variate relationships and temporal patterns. The dual masking strategy combines random masking (75% of cases) to learn bidirectional relationships and post-fix masking (25% of cases) to capture temporal causality. The NCC loss is added to the mean squared error (MSE) loss to capture global temporal patterns and long-range dependencies.
OTiS is pre-trained on a large and diverse corpus spanning 8 domains (ECG, Temperature, Audio, Electromechanics, EEG, Banking, Economics) with a total of 640,187 samples and 11 billion time points. The model is evaluated across 15 diverse benchmarks covering classification, regression, and forecasting tasks, and is compared against 27 baselines, including 15 specialised and 12 general models.
Key findings include:
-
OTiS achieves state-of-the-art performance on 10 out of 15 benchmarks, with competitive results on the remaining tasks.
-
In classification, OTiS shows particular strength in processing long time series, yielding large performance boosts on the FD-B dataset.
-
In regression, OTiS outperforms specialised models that are unimodally or even multimodally pre-trained.
-
In forecasting, OTiS effectively captures temporal patterns and trends, even in previously unseen domains.
-
Zero-shot experiments and linear probing demonstrate the generalisability of OTiS's extracted time series features.
The paper also includes a domain signature analysis, showing that the learned domain-specific variate embeddings accurately capture inter-variate relationships. For example, EEG-specific embeddings recover the 3D spatial arrangement of electrodes, ECG-specific embeddings reflect Einthoven's triangle and the 3D view of the heart, and Weather-specific embeddings capture true physical relationships among climatological indicators. Additionally, experiments on sine waves demonstrate that OTiS captures temporal patterns that generalise to unseen frequencies.
A scaling study shows that downstream performance generally scales with dataset size, while model size scaling requires even larger pre-training corpora to be effective. An ablation study confirms that all key components—domain-specific tokeniser, dual masking strategy, and NCC loss—are essential for optimal performance.
The paper concludes that OTiS establishes an essential foundation for future advancements in time series analysis, particularly valuable in domains with limited data availability and high privacy standards, such as medicine. Limitations include the reliance on manually curated data and the potential benefit of larger pre-training corpora.
Improvements for AI systems
Based on the paper, here are the specific improvements I can implement in AI systems:
Improvement: Replace uniform patch-based tokenization with a tokenizer that uses learnable domain-specific variate embeddings, which are added to patch and temporal embeddings before feeding into the transformer.
What the improved system can do:
-
Handle time series with varying numbers of variates (from 1 to 321+) without resampling or cropping that distorts data
-
Preserve inter-variate relationships (e.g., spatial arrangement of EEG electrodes, physical relationships between weather indicators)
-
Adapt to previously unseen domains by randomly initializing new variate embeddings and fine-tuning them with limited data
-
Process multi-variate time series jointly rather than treating each variate independently (unlike MOMENT or MOIRAI)
Abstract
Recent breakthroughs in natural language processing and computer vision, driven by efficient pre-training on large datasets, have enabled foundation models to excel on a wide range of tasks. However, this potential has not yet been fully realised in time series analysis, as existing methods fail to address the heterogeneity in large time series corpora. Prevalent in domains ranging from medicine to finance, time series vary substantially in characteristics such as variate count, inter-variate relationships, temporal patterns, and sampling frequency. To address this, we introduce a novel pre-training paradigm specifically designed to handle time series heterogeneity. We propose a tokeniser with learnable domain signatures, a dual masking strategy, and a normalised cross-correlation loss, enabling our open model for general time series analysis (OTiS) to efficiently learn from large time series corpora. Extensive benchmarking on diverse tasks, such as classification, regression, and forecasting, demonstrates that OTiS outperforms state-of-the-art baselines. Our code and pre-trained weights are available at https://github.com/oetu/otis.
Sources
- Ti-MAE: Self-Supervised Masked Time Series Autoencoders
- Scaling Laws for Neural Language Models
- Transformer-based Spatial-Temporal Feature Learning for EEG Decoding
- Exploring Contrastive Learning in Human Activity Recognition for Healthcare
- LLaMA: Open and Efficient Foundation Language Models
- SAM 2: Segment Anything in Images and Videos
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks