Towards Unified Approaches in Self-Supervised Event Stream Modeling: Progress and Prospects
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Towards Unified Approaches in Self-Supervised Event Stream Modeling: Progress and Prospects".
Jane: The paper was written by Levente Zólyomi, Tianze Wang, Sofiane Ennadir, Oleg Smirnov and Lele Cao from Johannes Kepler University and NXAI GmbH and Kreditz AB and King AI Labs and Microsoft Gaming.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title and Authors: Tom: Alright, welcome back to the show, everyone! We've got a fascinating new paper on the arXiv preprint server today, and it's called "Towards Unified Approaches in Self-Supervised Event Stream Modeling: Progress and Prospects." Jane, I have to say, just reading that title got me excited — it's a survey, but it's aiming to do something really ambitious.
Jane: It really is, Tom. And I love that it's coming from this international team — Levente Zólyomi from Johannes Kepler University and NXAI, Tianze Wang from Kreditz, Sofiane Ennadir from King AI Labs, and Oleg Smirnov and Lele Cao from Microsoft. That's a pretty diverse group, which makes sense for a paper trying to unify work across healthcare, finance, gaming, and e-commerce.
Tom: Right, and that's the key word — "unify." For years, researchers in each of those fields have been building self-supervised models for event streams, but they've been doing it in isolation. The healthcare people talk to healthcare people, the gaming people talk to gaming people, and nobody's really stepping back to say, "Hey, we're all solving the same fundamental problem here."
Jane: Exactly. And the fundamental problem is this: an event stream is just a sequence of timestamped events. A patient's hospital visits, a player's in-game actions, a customer's clicks and purchases — structurally, they're all the same kind of data. The paper makes that point really clearly in their Table two where they map the same concepts across all four domains.
Tom: So the authors are essentially saying, let's stop reinventing the wheel in each vertical and start sharing what works. That's a big deal for the field, because it could accelerate progress enormously. I mean, if you can take a pre-training method that works for healthcare data and apply it to gaming logs, you're saving months of research time.
Jane: And it's not just about saving time, Tom. It's about the fact that some domains have way more data than others. Healthcare has massive datasets like MIMIC, but they're hard to access due to privacy. Gaming has tons of data too, but it's mostly proprietary. If we can develop methods that work across domains, then a breakthrough in one field can lift up the others.
Tom: That's a beautiful way to put it. And I think the authors would agree — they explicitly say the field is "fragmented" and that this fragmentation is holding back progress. They're not just writing a summary of existing work; they're making an argument for how the field should organize itself going forward.
Jane: Right, and that's what makes this survey different from just a literature review. It's prescriptive, not just descriptive. They're saying, here's how we should think about all these methods, here's what's missing, and here's where we need to go. That's a roadmap, not just a map.
Tom: So before we get into the actual content of the survey, I want to flag something for our listeners — this paper has a really nice taxonomy of self-supervised learning methods, and it's going to be the backbone of our whole discussion today. Jane, you want to give us a quick preview?
Jane: Sure, Tom. They split everything into two big families: predictive methods, where the model learns by predicting missing or future parts of the stream, and contrastive methods, where the model learns by comparing different views of the same data. We'll dig into both in the next segment.
Tom: And I can already tell you, the way they organize this is going to make you see connections you never noticed before. Stick around, because we're just getting started with "Towards Unified Approaches in Self-Supervised Event Stream Modeling: Progress and Prospects."
Summary of the Paper: Jane: Welcome back, everyone. We're still on "Towards Unified Approaches in Self-Supervised Event Stream Modeling: Progress and Prospects," and now we need to get into the meat of it. Tom, you mentioned the taxonomy — let's break it down.
Tom: Please do, because I want to make sure our listeners really get this. The predictive side is the more established one, right?
Jane: Absolutely. Predictive methods are the workhorses of the field. The first big category is masked modeling — think BERT for event streams. You randomly hide some events or attributes, and the model has to guess what was hidden. BEHRT in healthcare does this with patient diagnoses, BERT4Rec does it with user-item interactions in e-commerce, and there's even BERT4Eth for blockchain fraud detection.
Tom: And then there's autoregressive modeling, which is more like GPT — you predict the next event given everything that came before. That's huge in sequential recommendation systems, where you're trying to guess what a user will buy next. The paper mentions a bunch of methods there, like the ones from Tang and Wang or Li and colleagues.
Jane: Right, and then there's a third predictive approach that's a bit different — temporal point processes, or TPPs. These are mathematical models that focus specifically on when events happen, not just what they are. They're great for things like high-frequency trading, where the timing of a transaction is just as important as the transaction itself.
Tom: So predictive methods are all about filling in the blanks, whether that's a missing event, the next event, or the timing of events. Got it. Now what about the contrastive side?
Jane: Contrastive methods are about learning what makes two event streams similar or different. The most common approach is instance contrastive — you take one stream, create two slightly different versions of it through augmentations, and teach the model to recognize that they're the same thing. CoLES does this for financial transactions, CL4SRec does it for recommendations.
Tom: And I remember from the paper that there are also distillation-based methods, where a student model learns to match a teacher model's representations, and feature decorrelation methods, which try to make sure different parts of the representation capture different information. Those are less common but really clever.
Jane: Exactly. And the paper makes a really interesting observation — contrastive methods are underutilized in event streams compared to predictive ones. They argue that's a missed opportunity, because contrastive learning works at the entity level — the patient, the player, the customer — rather than the event level. And since most downstream tasks care about entities, that could be a better fit.
Tom: That's a sharp insight. And it connects to something the authors say about noise — individual events can be noisy, like accidental clicks in e-commerce, but the overall pattern of an entity's behavior is more stable. Contrastive methods can capture that stability.
Jane: Right. And they also highlight that the field is heavily skewed — they reviewed around a hundred papers, and the vast majority use predictive methods. Only a handful use contrastive approaches, and even fewer use hybrid methods that combine both. That's a clear gap in the literature.
Tom: So the summary of the paper is basically: here's how the field is organized, here's what's working, and here's what's being neglected. And the neglected parts might actually be the most promising. That's a strong takeaway.
Jane: It is. And it sets up their suggestions for improvement, which we're going to talk about next. They have some really concrete ideas about where the field should go from here.
Tom: Can't wait. Before we move on, though, let me just remind our listeners — we're discussing "Towards Unified Approaches in Self-Supervised Event Stream Modeling: Progress and Prospects," and the next segment is where things get really exciting.
Improvements Suggested by the Paper: Tom: Alright, we're back with "Towards Unified Approaches in Self-Supervised Event Stream Modeling: Progress and Prospects," and now we get to the part I've been waiting for — what does this paper say we should actually do differently? Jane, what's the biggest improvement they're calling for?
Jane: The biggest one, Tom, is what they call "domain-agnostic learning." The paper argues that because event streams across healthcare, finance, gaming, and e-commerce share so much structural similarity, we should be building methods that work across all of them. Not just adapting a method from one domain to another, but designing methods from the ground up that don't care what domain they're in.
Tom: That's a bold vision. It's essentially saying we could have a foundation model for event streams, kind of like GPT is for text. And they're not the only ones thinking this way — the paper mentions that tools like Event Stream GPT are starting to emerge, though they're still limited to specific domains.
Jane: Right, and they're honest about the challenges. Event streams are heterogeneous — an event might have categorical codes, numerical measurements, and unstructured text all mixed together. Designing a single tokenization scheme that handles all of that is genuinely hard. And then there's the multi-entity problem — you have users, items, accounts, all interacting, and flattening that into a one-dimensional sequence loses information.
Lu: Can I jump in here? I've been listening and this is exactly where my mind goes. The paper's call for domain-agnostic models is great, but I think the deeper opportunity is in the contrastive methods they say are underused. If we can crack augmentation design for event streams — figuring out what transformations preserve meaning across domains — we could unlock representations that transfer across entirely different applications.
Jane: That's a great point, Lu. And the paper does highlight augmentation design as a critical open problem. In computer vision, you can rotate an image or change its colors and it's still the same object. But what's the equivalent for an event stream? Can you randomly drop events without losing meaning? Can you reorder them? The answers aren't obvious, and they probably depend on the domain.
Tom: And that's exactly why the paper also calls for better benchmarks. Right now, most evaluation happens within a single domain, on proprietary data, with inconsistent metrics. Healthcare uses AUROC, e-commerce uses hit rate and NDCG — you can't compare results across fields. The authors want standardized benchmarks that test methods across multiple domains.
Meng: As someone who actually has to deploy these models, I want to add something to that. The paper talks about the need for open datasets, and I think that's the single biggest bottleneck in practice. We can't release our gaming logs because they're business-sensitive. But if there were good synthetic benchmarks that captured realistic event stream properties — bursty timing, heavy-tailed event frequencies — we could validate methods without exposing proprietary data.
Jane: That's a really practical point, Meng. And the paper does mention synthetic data as a promising direction, especially for domains like cybersecurity and DevOps where real logs can't be shared. They even suggest that recent advances in synthetic data generation could help preserve data properties while protecting confidentiality.
Tom: So the improvements boil down to three things: build domain-agnostic methods, explore contrastive and hybrid paradigms more aggressively, and create shared benchmarks and datasets. That's a concrete agenda.
Lu: And I'd add one more from the paper — they want us to take timestamps seriously. Many methods either ignore timing entirely or treat it as just another feature. But the paper argues we need systematic studies of how to model time, including whether to mask timestamps during pre-training. That's an underappreciated design choice that could have big effects.
Jane: Exactly, Lu. And that's actually where we're heading next — the paper's vision for the future and what it means for the broader world. Tom, should we wrap this up?
Tom: Almost, Jane. But first, let me say — this paper isn't just a survey. It's a call to action. And I think our listeners are going to want to hear what that means for the future.
Conclusion: Tom: And we're back for the final segment on "Towards Unified Approaches in Self-Supervised Event Stream Modeling: Progress and Prospects." Jane, let's bring it home. What's the big picture here?
Jane: The big picture, Tom, is that event stream data is everywhere — every click, every purchase, every hospital visit, every in-game action — and we're only scratching the surface of what we can learn from it. Self-supervised learning lets us learn from all that unlabeled data without needing expensive human annotations. And this paper shows us how to do it better by learning across domains instead of within them.
Lu: I want to build on that. The paper's vision of domain-agnostic SSL for event streams could genuinely transform industries. In healthcare, it could mean more accurate risk prediction with less labeled data, which matters because labels in medicine are expensive and slow to produce. In e-commerce, it could mean recommender systems that understand users more deeply. The same core technology, applied across all of them.
Meng: And from my side, the practical impact is about scalability. If we can build one model that handles multiple domains, we don't need to train and maintain separate systems for each product. That's a huge efficiency gain. But it only works if we solve the benchmark problem — we need to be able to measure whether a domain-agnostic model actually performs well everywhere.
Jane: That's exactly why the paper's call for open benchmarks is so important. Without them, we can't know if a method that works for healthcare will work for gaming. We're flying blind.
Tom: And the authors are also pushing us to think about the ethical dimension. Event streams often contain sensitive information — patient records, financial transactions, player behavior. As we build more powerful models, we need to think about privacy, fairness, and transparency. The paper mentions federated learning as one way to train across institutions without sharing raw data.
Lu: Right, and that's not an afterthought. If we're building foundation models for event streams, we're building systems that will make decisions about people — whether they get a loan, whether they need medical intervention, what content they see. We have to get this right.
Jane: Absolutely. So let me summarize what we learned today. "Towards Unified Approaches in Self-Supervised Event Stream Modeling: Progress and Prospects" gives us a taxonomy of self-supervised methods — predictive and contrastive — shows us that predictive methods dominate but contrastive methods hold untapped potential, and lays out a roadmap for the field: build domain-agnostic models, explore new paradigms, create shared benchmarks, and take timestamps seriously.
Tom: And the takeaway for our listeners — whether you're a researcher, an engineer, or just someone curious about AI — is that event streams are a unifying language across industries, and self-supervised learning is the key to understanding them. This paper is a milestone in making that vision concrete.
Jane: Well said, Tom. We're going to say goodbye to this paper now, but I have a feeling we'll be seeing its influence for years to come.
Tom: Same here, Jane. Thanks to Lu and Meng for joining us today — great insights as always. And to our listeners, stay curious, keep learning, and we'll see you on the next episode.
Jane: Goodbye, everyone!
Johannes Kepler University · NXAI GmbH · Kreditz AB · King AI Labs · Microsoft Gaming
cs.LG, cs.AI
Submitted: 2025-02-07
Updated: 2026-09-11
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
Importance score: 57/100
The gist: This survey paper, "Towards Unified Approaches in Self-Supervised Event Stream Modeling: Progress and Prospects" by Levente Zólyomi, Tianze Wang, Sofiane Ennadir, Oleg Smirnov, and Lele Cao,
Key concepts
- Self-Supervised Event Stream Modeling
- This involves using models trained on sequences of timestamped events—like patient visits or user clicks—to learn patterns without needing human labels. The goal is to find methods that work across different data types, such as those found in healthcare or e-commerce.
- Predictive Methods
- These are established modeling techniques where the system learns by predicting missing parts of a stream. Examples include masked modeling (like BERT, guessing hidden events) or autoregressive modeling (predict-the-next event), focusing on filling in blanks within the sequence.
- Contrastive Methods
- This approach teaches a model to determine if two event streams are similar or different. It often works by creating slightly varied versions of one stream and then comparing them, helping capture stable patterns at the entity level (e,g., a specific user).
- Domain-Agnostic Learning
- The paper advocates for building models that do not care about the specific field they come from—such as finance or gaming. Instead of adapting existing methods, this design methods from the ground up to handle the structural similarity across all event streams.
Terminology
Summary
This survey paper, Towards Unified Approaches in Self-Supervised Event Stream Modeling: Progress and Prospects
by Levente Zólyomi, Tianze Wang, Sofiane Ennadir, Oleg Smirnov, and Lele Cao, systematically reviews and synthesizes Self-Supervised Learning (SSL) methodologies tailored for event stream (ES) modeling across multiple domains.
The paper addresses the challenge that "large-scale event datasets typically lack the extensive labeling needed to train traditional supervised learning systems. Moreover, existing modeling efforts often remain fragmented across domains, sometimes replicating research efforts without leveraging the many structural similarities shared by event streams across industries. The authors note that
Self-Supervised Learning (SSL) has emerged as a promising paradigm to address these challenges by enabling the extraction of meaningful representations from unlabeled ES data."
The paper formally defines an event stream: An ES is a continuous, ordered sequence of events generated over time by one or more sources. Each event encapsulates timestamped information about a specific action or state.
Formally, an ES for an entity u can be defined as a sequence: Su = eu,i ∞ i=1, where eu,i = (tu,i, du,i)
where tu,i ∈ T is the timestamp of eu,i from a totally ordered time domain T
and du,i ∈ D represents all information associated with eu,i.
The survey offers several key contributions: (1) Comprehensive synthesis across domains: We provide the first (to our knowledge) cross-domain assessment of SSL-based ES modeling
; (2) Taxonomy of SSL paradigms for ES: We present a structured taxonomy that organizes the wide spectrum of SSL methods used in ES modeling, ranging from predictive techniques to various contrastive approaches
; (3) Resource and benchmark guide: We compile an overview of widely used public datasets, identify limitations in current benchmarking practices
; and (4) Roadmap for future research: We outline critical open issues in ES SSL.
The paper organizes SSL methods into two major families:
Masked Modeling: In masked (or denoising) modeling, the learner randomly masks out certain events, event attributes, or timestamps within a sequence and aims to predict the masked portions using the remaining context.
Examples include BEHRT, BRLTM, and MedBERT in healthcare; BERT4Rec in e-commerce; BERT4Eth in finance; and Pu et al. (2022) and Player2Vec in gaming.
Autoregressive Modeling: Autoregressive (AR) approaches for ES revolve around predicting the next event (or next attribute) given the historical context.
These methods extend concepts from language modeling, forcing a strictly causal order – using only past events to predict future ones.
Many sequential recommender systems adopt this perspective.
Temporal Point Processes (TPPs): "TPPs are a generative modeling framework designed to specify the probability distribution over when events occur. Rather than predicting the next token in a sequence, TPPs specify a conditional intensity function that depends on the history of previous events to determine the timing of future ones. The paper notes that
TPPs occupy a distinct position in this taxonomy and that
their inclusion under predictive SSL reflects both practical importance and conceptual overlap."
Instance Contrastive: "Instance-based contrastive learning aims to learn discriminative representations by treating different augmentations of the same event stream instance as positive pairs and instances (or their augmentations) from other entities as negative pairs." Examples include CL4SRec in sequential recommenders, CoLES in finance, and Raghu et al. (2023) in healthcare.
Distillation: Distillation-based contrastive learning avoids explicit negative samples by coupling a teacher network and a student network on differently augmented views of the same event stream.
The paper notes Hi-BEHRT as a notable example in healthcare.
Feature Decorrelation: Feature decorrelation methods aim to learn representations with minimal redundancy across latent dimensions, typically without explicit negative pairs.
Raghu et al. (2023) is cited as adopting a VICReg-type loss for healthcare EHR data.
Multimodal Contrastive: Multimodal contrastive learning aims to align representations across multiple modalities associated with the same data instance.
Examples include King et al. (2023) and Ma et al. (2024) in healthcare settings.
The paper covers four primary domains: healthcare (EHRs, patient trajectories), e-commerce (sequential recommender systems), finance (fraud detection, transaction modeling), and gaming (player behavior modeling). The authors note that "ES data in these domains shares common difficulties: (1) irregular sampling, as event arrivals can be sporadic, requiring flexible timestamp modeling; (2) diverse feature types, because each event can encompass a wide range of features, from discrete labels to unstructured text or images; and (3) privacy and ethical concerns."
The paper reviews commonly used datasets including CPRD, MIMIC, Amazon, Yelp, MovieLens, StackOverflow, and Retweet. The authors note that "Public data availability varies considerably across domains: healthcare datasets represent carefully de-identified releases from academic institutions, while e-commerce datasets primarily consist of reviews and ratings rather than granular behavioral logs. Gaming and finance are notably underrepresented."
The paper identifies several critical gaps and future directions:
-
Domain-agnostic Learning:
The consistent structure of ES data and the alignment of learning objectives across domains present an opportunity to design generalized SSL frameworks for ES modeling.
-
Underexplored Modern Paradigms: "Research in SSL has made significant progress across various domains, yet the majority of methods in ES remain focused on predictive learning objectives... This narrow focus leaves other promising directions, such as contrastive learning, relatively unexplored."
-
Event Time Modeling:
The role of timestamps in self-supervised ES modeling remains underexplored, with limited systematic evaluation of their influence on model performance.
-
Open Datasets and Benchmarks:
The scarcity of datasets is likely due to the nature of the data, as it is often generated by human actions... is often considered business-sensitive or private.
-
Large Language and Foundation Models: "Most existing methods can be viewed as using 'LLM-style' architectures as building blocks for self-supervised pre-training, but without yet reaching the scale, cross-domain coverage, or emergent capabilities that are typically associated with foundation models in other modalities."
The paper concludes that SSL has emerged as a promising approach to address the inherent challenges of modeling ES data across diverse domains
but the field remains fragmented, with predictive SSL paradigms dominating research efforts and other methodologies, such as contrastive learning, underexplored.
The authors state that future research should focus on developing domain-agnostic ES models, systematically integrating event timing, and exploring innovative paradigms such as contrastive, joint-embedding predictive, and hybrid architectures.
Improvements for AI systems
Based on the paper, here are the specific improvements I can make to AI systems and what the improved systems can do:
Improvement: Build a single pre-trained model that learns representations from unlabeled event streams across healthcare, e-commerce, finance, and gaming simultaneously, using a shared tokenization scheme for heterogeneous event attributes (categorical codes, numerical values, timestamps).
What the improved system can do: A model pre-trained on diverse event streams can transfer to a new domain with minimal fine-tuning—e.g., a model pre-trained on e-commerce clicks and EHR codes can be adapted to detect fraudulent financial transactions or predict player churn with only a few hundred labeled examples, rather than requiring domain-specific training from scratch.
Abstract
The proliferation of digital interactions across diverse domains, such as healthcare, e-commerce, gaming, and finance, has resulted in the generation of vast volumes of event stream (ES) data. ES data comprises continuous sequences of timestamped events that encapsulate detailed contextual information relevant to each domain. While ES data holds significant potential for extracting actionable insights and enhancing decision-making, its effective utilization is hindered by challenges such as the scarcity of labeled data and the fragmented nature of existing research efforts. Self-Supervised Learning (SSL) has emerged as a promising paradigm to address these challenges by enabling the extraction of meaningful representations from unlabeled ES data. In this survey, we systematically review and synthesize SSL methodologies tailored for ES modeling across multiple domains, bridging the gaps between domain-specific approaches that have traditionally operated in isolation. We present a comprehensive taxonomy of SSL techniques, encompassing both predictive and contrastive paradigms, and analyze their applicability and effectiveness within different application contexts. Furthermore, we identify critical gaps in current research and propose a future research agenda aimed at developing scalable, domain-agnostic SSL frameworks for ES modeling. By unifying disparate research efforts and highlighting cross-domain synergies, this survey aims to accelerate innovation, improve reproducibility, and expand the applicability of SSL to diverse real-world ES challenges.
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks