PRAGMA: Revolut Foundation Model

summary

Video file (mp4)

The gist

PRAGMA is introduced as a family of foundation models designed for multi-source banking event sequences, addressing the challenge that "modern financial systems generate vast quantities of

In short

PRAGMA is a foundation model designed to unify complex, variable-length user histories into embeddings. It uses a key–value–time tokenization and dual encoder branches to capture both static profile state and dynamic event sequences. The model achieves superior performance across six downstream tasks, including fraud detection, and is highly efficient using LoRA fine-tuning for deployment.

Key concepts

Key–Value–Time Tokenization
This is a novel tokenization scheme used to handle varied data types within user histories. It allows the model to treat three distinct components—the semantic meaning of a field, its actual value, and its timing—as separate entities, providing an elegant solution for mapping unstructured reality onto the model's rigid structure.
Dual Temporal Encoding
PRAGMA encodes time in two ways: by calculating the log-seconds since the last event and embedding cyclical features like the day of the week. This dual approach captures both precise sequence timing for recent actions and the predictable daily rhythm of user behavior, improving efficiency.
Encoder-Only Bidirectional Design
PRAGMA utilizes an encoder-only, bidirectional architecture. The primary goal is to create robust, transferable representations of data structure rather than generating new text. This allows the model to learn the underlying structure of a user's life.
LoRA Fine-tuning
LoRA (Low-Rank Adaptation) is a practical method used for fine-tuning the large foundation model. The paper notes that using this technique can match or exceed training from scratch, making it highly efficient for managing the operational costs of deploying such a massive system.

Terminology used across episodes

This episode discusses

The paper

PRAGMA: Revolut Foundation Model · Read on arXiv

Xiyuan Zhang, Danielle C. Maddix, Junming Yin, Nick Erickson, Abdul Fatir Ansari, Boran Han, Shuai Zhang, Leman Akoglu, Christos Faloutsos, Michael W. Mahoney, Cuixiong Hu, Huzefa Rangwala, George Karypis, and Bernie Wang

Modern financial systems generate vast quantities of transactional and event-level data that encode rich economic signals. This paper presents PRAGMA, a family of foundation models for banking event sequences. Our approach pre-trains a Transformer-based architecture with masked modelling on a large-scale, heterogeneous banking event corpus using a self-supervised objective tailored to the discrete, variable-length nature of financial records. The resulting model supports a wide range of downstream tasks such as credit scoring, fraud detection, and lifetime value prediction: strong performance can be achieved by training a simple linear model on top of the extracted embeddings and can be further improved with lightweight fine-tuning. Through extensive evaluation on downstream tasks, we demonstrate that PRAGMA achieves superior performance across multiple domains directly from raw event sequences, providing a general-purpose representation layer for financial applications.

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "PRAGMA: Revolut Foundation Model".

Jane: The paper was written by Xiyuan Zhang, Danielle C. Maddix, Junming Yin, Nick Erickson, Abdul Fatir Ansari et al. from.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Paper discussion segment 2: Jane: The paper summarizes its approach by tackling the inherent difficulties in modeling complex user histories, which are not like standard text because they're variable-length records containing mixed types of data.

Tom: And it seems like they’ve found that just serializing these structured records as text and feeding them into a basic Transformer is too inefficient because of all the field names and delimiters.

Lu: That complexity is compounded by how the sequential nature of user activity—the long-tailed histories—makes traditional tokenization struggle to capture the true temporal relationships between actions.

Meng: The paper describes PRAGMA as an encoder-only, bidirectional design, which makes sense because their primary goal is to create robust, transferable representations rather than trying to generate new text.

Lalam: I like that focus on representation; it means the model learns the structure of the user's life, not just a sequence of words.

Tom: The paper mentions a key innovation in this summary: combining multisource events with static profile state, which is critical for making sure you have all the context needed to make a prediction.

Jane: It also highlights how they are building this foundation model by masking parts of the data and trying to reconstruct them, which is essentially their core training objective.

Lu: This masked modeling approach ensures that the model learns deep, structural understanding from observing partial histories, which is far more robust than simple pattern matching.

Meng: It’s a clever way to force the the AI to learn context from missing data points in a real-world transaction sequence.

Lalam: And by focusing on this representation layer, we are moving away from building dozens of small, task-specific pipelines toward a single general solution.

Paper discussion segment 3: Tom: Now, let's talk about the actual technical improvements—how does PRAGMA achieve its superior performance? The paper details some very specific engineering choices to handle this heterogeneity.

Jane: They've moved beyond just using a standard tokenization scheme by introducing a key–value–time tokenization approach for handling the varied data types.

Lu: This allows us to treat the semantic meaning of a field as one thing, its actual value as another, and its timing as yet another, which is such an elegant solution to mapping unstructured reality onto a rigid model.

Meng: The paper explains that they use two separate encoder branches—one for the static profile state and one for the event sequence—which feed into a final history encoder.

Lalam: This dual-branch design is powerful because it ensures the model always has access to both who the user is (the profile) and what they did (the events).

Tom: It seems like this whole structure, combined with their specific method for handling time, really sets them apart from previous transaction-ledger models.

Jane: The paper describes how they encode time in two ways: calculating the log-seconds since the last event, plus embedding cyclical features like day of week.

Lu: That dual temporal encoding is genius because it captures both the precise sequence timing for recent events and the predictable daily rhythm of user behavior.

Meng: From an implementation view, using this approach allows for much more efficient training than trying to force all raw data into a single, giant string.

Lalam: When we are able to process both the static traits and the dynamic actions through structured embeddings, we are drastically improving how AI can understand complex financial context.

Paper discussion segment 4: Tom: The paper then moves into the results, showing how PRAGMA performs against various internal benchmarks—it’s not just one task, it's a whole suite.

Jane: The performance is consistently superior across six different downstream tasks, including credit scoring and fraud detection.

Lu: But what stands out is the scale of the performance boost in certain areas, like the one hundred thirty point two percent relative increase in PR-AUC for communication engagement.

Meng: That’s a huge gain, and the paper shows that this performance scales up when they move from their Small version to their Large one billion parameter model.

Lalam: The idea of finding a general representation that works well across all six tasks is what I find most impressive; the AI is learning universal financial patterns.

Tom: It's clear they are reducing the need for those massive, hand-crafted features that traditionally required specialized data science teams for one task.

Jane: The paper suggests that even a small model, PRAGMA-S, can be highly competitive, which is great news for efficiency and scaling.

Lu: It also notes that in some tasks like Lifetime Value, the gains are more modest because those behaviors might already be well-captured by existing models.

Meng: That makes sense—if the data is simple enough, we don're not going to need a one billion parameter model to solve it.

Lalam: The consistent performance across all six diverse benchmarks suggests that this general representation layer is robust and reliable for AI in finance.

Conclusion: Tom: So, after looking at the architecture, the methods, and the results, we're left with a powerful new tool for understanding financial data.

Jane: The paper concludes that PRAGMA successfully provides a general-purpose representation layer by unifying user histories into embeddings.

Lu: I think this is a major step toward treating raw banking ledgers as something comparable to massive text corpora, enabling the Transformer architecture to shine in finance.

Meng: It’s also very practical that they found LoRA fine-tuning can match or exceed training from scratch, which is huge for managing the operational costs of deploying such a large system.

Lalam: My takeaway is that this confirms that complex financial behavior—the way users interact over time—can be modeled with the same level of depth as other data streams, which is great for AI culture overall.

Tom: We've seen how PRAGMA works, from its key-value-time tokenization to the impressive gains in credit scoring and fraud detection.

Jane: It’s clear that this represents a major shift in how we approach consumer banking data analysis.

Lu: I hope future work on cross-record dependencies can address those relational challenges they noted at the end.

Meng: I'm looking forward to seeing how this scales into production environments, given its efficiency gains with LoRA.

Lalam: We're really excited to see the impact of PRAGMA on how AI understands and supports financial services going forward.

Tom: That’s a great summary, Jane; it’s definitely a paradigm worth keeping an eye on.

Jane: Absolutely, Tom; we'll be looking for that "PRAGMA: Revolut Foundation Model" in the next paper to wrap up our discussion today.

More episodes

← Home