PRAGMA: Revolut Foundation Model
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "PRAGMA: Revolut Foundation Model".
Jane: The paper was written by Xiyuan Zhang, Danielle C. Maddix, Junming Yin, Nick Erickson, Abdul Fatir Ansari et al. from.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Paper discussion segment 2: Jane: The paper summarizes its approach by tackling the inherent difficulties in modeling complex user histories, which are not like standard text because they're variable-length records containing mixed types of data.
Tom: And it seems like they’ve found that just serializing these structured records as text and feeding them into a basic Transformer is too inefficient because of all the field names and delimiters.
Lu: That complexity is compounded by how the sequential nature of user activity—the long-tailed histories—makes traditional tokenization struggle to capture the true temporal relationships between actions.
Meng: The paper describes PRAGMA as an encoder-only, bidirectional design, which makes sense because their primary goal is to create robust, transferable representations rather than trying to generate new text.
Lalam: I like that focus on representation; it means the model learns the structure of the user's life, not just a sequence of words.
Tom: The paper mentions a key innovation in this summary: combining multisource events with static profile state, which is critical for making sure you have all the context needed to make a prediction.
Jane: It also highlights how they are building this foundation model by masking parts of the data and trying to reconstruct them, which is essentially their core training objective.
Lu: This masked modeling approach ensures that the model learns deep, structural understanding from observing partial histories, which is far more robust than simple pattern matching.
Meng: It’s a clever way to force the the AI to learn context from missing data points in a real-world transaction sequence.
Lalam: And by focusing on this representation layer, we are moving away from building dozens of small, task-specific pipelines toward a single general solution.
Paper discussion segment 3: Tom: Now, let's talk about the actual technical improvements—how does PRAGMA achieve its superior performance? The paper details some very specific engineering choices to handle this heterogeneity.
Jane: They've moved beyond just using a standard tokenization scheme by introducing a key–value–time tokenization approach for handling the varied data types.
Lu: This allows us to treat the semantic meaning of a field as one thing, its actual value as another, and its timing as yet another, which is such an elegant solution to mapping unstructured reality onto a rigid model.
Meng: The paper explains that they use two separate encoder branches—one for the static profile state and one for the event sequence—which feed into a final history encoder.
Lalam: This dual-branch design is powerful because it ensures the model always has access to both who the user is (the profile) and what they did (the events).
Tom: It seems like this whole structure, combined with their specific method for handling time, really sets them apart from previous transaction-ledger models.
Jane: The paper describes how they encode time in two ways: calculating the log-seconds since the last event, plus embedding cyclical features like day of week.
Lu: That dual temporal encoding is genius because it captures both the precise sequence timing for recent events and the predictable daily rhythm of user behavior.
Meng: From an implementation view, using this approach allows for much more efficient training than trying to force all raw data into a single, giant string.
Lalam: When we are able to process both the static traits and the dynamic actions through structured embeddings, we are drastically improving how AI can understand complex financial context.
Paper discussion segment 4: Tom: The paper then moves into the results, showing how PRAGMA performs against various internal benchmarks—it’s not just one task, it's a whole suite.
Jane: The performance is consistently superior across six different downstream tasks, including credit scoring and fraud detection.
Lu: But what stands out is the scale of the performance boost in certain areas, like the one hundred thirty point two percent relative increase in PR-AUC for communication engagement.
Meng: That’s a huge gain, and the paper shows that this performance scales up when they move from their Small version to their Large one billion parameter model.
Lalam: The idea of finding a general representation that works well across all six tasks is what I find most impressive; the AI is learning universal financial patterns.
Tom: It's clear they are reducing the need for those massive, hand-crafted features that traditionally required specialized data science teams for one task.
Jane: The paper suggests that even a small model, PRAGMA-S, can be highly competitive, which is great news for efficiency and scaling.
Lu: It also notes that in some tasks like Lifetime Value, the gains are more modest because those behaviors might already be well-captured by existing models.
Meng: That makes sense—if the data is simple enough, we don're not going to need a one billion parameter model to solve it.
Lalam: The consistent performance across all six diverse benchmarks suggests that this general representation layer is robust and reliable for AI in finance.
Conclusion: Tom: So, after looking at the architecture, the methods, and the results, we're left with a powerful new tool for understanding financial data.
Jane: The paper concludes that PRAGMA successfully provides a general-purpose representation layer by unifying user histories into embeddings.
Lu: I think this is a major step toward treating raw banking ledgers as something comparable to massive text corpora, enabling the Transformer architecture to shine in finance.
Meng: It’s also very practical that they found LoRA fine-tuning can match or exceed training from scratch, which is huge for managing the operational costs of deploying such a large system.
Lalam: My takeaway is that this confirms that complex financial behavior—the way users interact over time—can be modeled with the same level of depth as other data streams, which is great for AI culture overall.
Tom: We've seen how PRAGMA works, from its key-value-time tokenization to the impressive gains in credit scoring and fraud detection.
Jane: It’s clear that this represents a major shift in how we approach consumer banking data analysis.
Lu: I hope future work on cross-record dependencies can address those relational challenges they noted at the end.
Meng: I'm looking forward to seeing how this scales into production environments, given its efficiency gains with LoRA.
Lalam: We're really excited to see the impact of PRAGMA on how AI understands and supports financial services going forward.
Tom: That’s a great summary, Jane; it’s definitely a paradigm worth keeping an eye on.
Jane: Absolutely, Tom; we'll be looking for that "PRAGMA: Revolut Foundation Model" in the next paper to wrap up our discussion today.
Xiyuan Zhang, Danielle C. Maddix, Junming Yin, Nick Erickson, Abdul Fatir Ansari, Boran Han, Shuai Zhang, Leman Akoglu, Christos Faloutsos, Michael W. Mahoney, Cuixiong Hu, Huzefa Rangwala, George Karypis, and Bernie Wang
cs.LG, cs.CE, cs.CL, cs.IR, q-fin.CP
Submitted: 2026-08-24
Updated: 2026-08-25
Comments: [v2]: adds extra ablations and results; related work improvements
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 81/100
The gist: PRAGMA is introduced as a family of foundation models designed for multi-source banking event sequences, addressing the challenge that "modern financial systems generate vast quantities of
Key concepts
- Key–Value–Time Tokenization
- This is a novel tokenization scheme used to handle varied data types within user histories. It allows the model to treat three distinct components—the semantic meaning of a field, its actual value, and its timing—as separate entities, providing an elegant solution for mapping unstructured reality onto the model's rigid structure.
- Dual Temporal Encoding
- PRAGMA encodes time in two ways: by calculating the log-seconds since the last event and embedding cyclical features like the day of the week. This dual approach captures both precise sequence timing for recent actions and the predictable daily rhythm of user behavior, improving efficiency.
- Encoder-Only Bidirectional Design
- PRAGMA utilizes an encoder-only, bidirectional architecture. The primary goal is to create robust, transferable representations of data structure rather than generating new text. This allows the model to learn the underlying structure of a user's life.
- LoRA Fine-tuning
- LoRA (Low-Rank Adaptation) is a practical method used for fine-tuning the large foundation model. The paper notes that using this technique can match or exceed training from scratch, making it highly efficient for managing the operational costs of deploying such a massive system.
Terminology
Summary
PRAGMA is introduced as a family of foundation models designed for multi-source banking event sequences, addressing the challenge that modern financial systems generate vast quantities of transactional and event-level data that encode rich economic signals
which are difficult to model efficiently using standard text tokenization.
The paper details PRAGMA's architecture and methodology:
** Architecture and Tokenization:** The model is an encoder-only, bidirectional design, consisting of three main blocks: the Profile State Encoder, the Event Encoder, and the History Encoder. To handle heterogeneity in financial records—which include mixed categorical, numerical, and free-text fields—PRAGMA employs a key–value–time tokenization scheme with type-specific value encoding.
The input is decomposed into a semantic type (key), a value (tokenized based on its type: percentile buckets for numbers, single tokens for categories, or subword tokens for text), and a temporal coordinate.
** Pre-training:** The model is pre-trained using a masked modelling (MLM) objective on a large-scale, heterogeneous banking event corpus.
This process allows the the model to learn record-level representations from complete histories
by masking individual tokens, entire events, or semantic types.
** Downstream Adaptation:** After pre-training, PRAGMA supports downstream tasks through two methods:
-
Embedding Probe: The frozen History Encoder output is used as a feature vector, and a lightweight linear model is trained on top.
-
LoRA Fine-tuning: Low-Rank Adaptation (LoRA) is applied to update only a small fraction of parameters, enabling
fast specialisation while keeping most of the backbone shared across tasks.
** Evaluation and Performance:** The authors evaluate PRAGMA on a suite of internal benchmarks, including credit scoring, fraud detection, communication engagement, recurrent transaction detection, lifetime value prediction (LTV), and product recommendation. The results demonstrate that PRAGMA consistently outperforms strong task-specific baselines while reducing the need for hand-crafted features.
** Key Findings:**
-
Model Scale: PRAGMA is available in three variants: 10 M, 100 M, and 1 B parameters. Scaling yields performance gains that are highly task-dependent; for example, the Large (1 B) model achieves a significant boost in PR-AUC for Credit Scoring.
-
LoRA Efficacy: The results validate that
LoRA fine-tuning consistently matches or exceeds the performance of full training from scratch across all evaluated tasks.
-
Profile State Value: The dedicated Profile State Encoder is particularly effective for tasks where static contextual attributes are informative, such as credit scoring, providing a
significant value
that event sequences alone cannot capture. -
** Limitations:** A critical limitation noted in the Anti-Money Laundering (AML) task is that PRAGMA
suffers a 47.1 % drop
compared to the production baseline, attributed to the fact thatPRAGMA processes event histories in isolation
and does not inherently capture cross-record dependency structures crucial for relational tasks.
In conclusion, PRAGMA provides a general-purpose representation layer for financial applications,
allowing a single pre-trained backbone to achieve superior performance across diverse downstream tasks.
Improvements for AI systems
Based on a rigorous analysis of the PRAGMA framework, I have identified several critical areas for architectural and methodological enhancement. These improvements address current limitations—particularly its inability to model cross-record dependencies—and optimize its utility for high-stakes commercial deployment.
The core limitation of PRAGMA is that it processes event histories in isolation, making it fundamentally unable to capture network-level or cross-record dependencies (as demonstrated by its poor performance on Anti-Money Laundering).
Improvement: Implement a Multi-User/Graph Aggregation Layer. Instead of processing individual user histories (z = [z a: z e]), the system must be adapted to process batches of users or even entire sub-networks. This requires integrating a Graph Neural Network (GNN) layer—such as Graph Attention Network (GAT)—after the History Encoder.
-
Mechanism: The output embeddings (z h) for all users in a batch are passed through GAT layers, where the adjacency matrix is derived from known relationships (e.g., shared IP addresses, linked accounts, or common merchants).
-
Function: This allows the model to learn how an individual's activity relates to the broader network structure, transforming PRAGMA from a sequence model into a relational inference engine.
While PRAGMA uses RoPE (Rotary Positional Embeddings) and log-seconds encoding, this is insufficient for capturing complex interactions spanning months or years.
-
Mechanism: Implement a secondary, slower recurrent module (e.g., a GRU or specialized memory cell) that accumulates state over large time intervals (e.g., monthly summaries), and then use attention mechanisms to inject this compressed long-term state into the current event encoding process.
-
Function: This enables the model to retain
memory
of milestones (like account age or past high-value transactions) without requiring a single, massive sequence length, solving the problem of forgetting early historical signals.
The current approach uses a custom BPE tokenizer and trainable embeddings for text fields. While effective for structured data, this misses the vast semantic knowledge embedded in large language models (LLMs).
-
Mechanism: The output of the frozen LLM is projected into the PRAGMA's embedding dimension (d model), and this semantic vector is fused with the standard token embeddings via a weighted sum before entering the Event Encoder.
-
Function: This allows PRAGMA to leverage external, high-quality linguistic knowledge for unstructured data (like transaction descriptions or notes), improving pattern recognition in text-heavy domains (e.g., fraud detection).
The current training and architecture is optimized for research performance, not necessarily for low-latency inference in a production environment.
-
Mechanism: Use a teacher-student training paradigm. Implement hardware acceleration layers (e.g., Triton kernels) specifically designed for the PRAGMA's unique two-branch/history encoder structure.
-
Function: This drastically reduces computational overhead and memory footprint, enabling real-time inference crucial for applications like fraud detection where latency must be in milliseconds, while maintaining the high accuracy of the large model.
By implementing these enhancements, PRAGMA evolves from a powerful sequence analyzer into a comprehensive Enterprise Risk and Predictive Intelligence Platform. The improved system will achieve:
-
Holistic Risk Assessment (AML/Regulatory Compliance): It can detect highly complex, distributed fraud or money laundering schemes by analyzing the relationship between multiple user accounts and transactions across different entities, providing actionable alerts where traditional models fail.
-
Hyper-Personalized Customer Journey Optimization: By combining sequential event data with long-term memory and static profile state, it can predict not just if a user will engage (e.g., open a re-engagement communication), but when and which specific product they are most likely to adopt next, enabling truly context-aware marketing.
-
Superior Predictive Accuracy in Niche Domains: It provides high-impact insights into complex financial metrics (like predicting gross profit or identifying subscription recurrence) by consolidating signals that were previously siloed in disparate feature engineering pipelines.
-
Operational Scalability: It allows for the deployment of state-of-the-art, highly accurate models at a fraction of the computational cost, making advanced AI accessible and affordable for continuous, real-time monitoring in production environments.
Abstract
Modern financial systems generate vast quantities of transactional and event-level data that encode rich economic signals. This paper presents PRAGMA, a family of foundation models for banking event sequences. Our approach pre-trains a Transformer-based architecture with masked modelling on a large-scale, heterogeneous banking event corpus using a self-supervised objective tailored to the discrete, variable-length nature of financial records. The resulting model supports a wide range of downstream tasks such as credit scoring, fraud detection, and lifetime value prediction: strong performance can be achieved by training a simple linear model on top of the extracted embeddings and can be further improved with lightweight fine-tuning. Through extensive evaluation on downstream tasks, we demonstrate that PRAGMA achieves superior performance across multiple domains directly from raw event sequences, providing a general-purpose representation layer for financial applications.
Sources
- BEiT: BERT Pre-Training of Image Transformers
- Your Spending Needs Attention: Modeling Financial Habits with Transformers
- NV-Retriever: Improving text embedding models with effective hard-negative mining
- TransactionGPT
- Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
- TabPFN-2.5: Advancing the State of the Art in Tabular Foundation Models
- Gaussian Error Linear Units (GELUs)
- TabTransformer: Tabular Data Modeling Using Contextual Embeddings
- GPT-4o System Card
- Mixtral of Experts
- Scaling Recommender Transformers to One Billion Parameters
- Muon is Scalable for LLM Training
- RoBERTa: A Robustly Optimized BERT Pretraining Approach
- SAINT: Improved Neural Networks for Tabular Data via Row Attention and Contrastive Pre-Training
- BloombergGPT: A Large Language Model for Finance
- TransAct V2: Lifelong User Action Sequence Modeling on Pinterest Recommendation
- FinBERT: A Pretrained Language Model for Financial Communications
- Actions Speak Louder than Words: Trillion-Parameter Sequential Transducers for Generative Recommendations
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks