Maxim Ostroukhov, Ruslan Mikhailov, Vladimir Iashin, Artem Sokolov, Andrei Akshonov, Vitaly Protasov, Andrey Goncharov, Dmitrii Beloborodov, Vince Mullin, Roman Y. Enzmann, Georgios Kolovos, Jason Renders, Pavel Nesterov, Anton Repushko
4 min
Modern financial institutions generate massive, heterogeneous event streams—ranging from card transactions to in-app navigation—that are difficult to model with standard language-model architectures. The authors investigate whether a single, large-scale foundation model can learn universal representations of these banking user histories that transfer effectively across multiple, disparate downstream financial tasks.
The researchers introduce PRAGMA, an encoder-only Transformer architecture designed specifically for multi-source banking data. Unlike standard LLMs that treat all data as text, PRAGMA uses a specialized tokenization scheme that preserves the structure of tabular financial records by decomposing them into semantic keys, values (numerical, categorical, or textual), and temporal coordinates. The model employs a two-branch design: a Profile State Encoder for static user attributes and an Event Encoder for sequential interactions, which are then fused by a History Encoder. The model is pre-trained using a masked modeling objective on a large-scale corpus of 24 billion events. For downstream tasks, the authors utilize either frozen embedding probes or parameter-efficient fine-tuning via Low-Rank Adaptation (LoRA).
PRAGMA demonstrates superior performance across six diverse downstream tasks, including credit scoring, fraud detection, and product recommendation, consistently outperforming strong, task-specific baselines. The authors show that scaling the model from 10 million to 1 billion parameters yields significant performance gains, particularly in high-impact tasks like credit scoring. Furthermore, LoRA fine-tuning of the pre-trained backbone consistently matches or exceeds the performance of training models from scratch, confirming the effectiveness of the learned representations. The architecture also proves robust to event staleness, making it suitable for real-time applications.
This work establishes that banking event sequences, despite their irregular timing and heterogeneous structure, can benefit from the foundation model paradigm. By consolidating multiple independent, high-maintenance pipelines into a single shared backbone, financial institutions can reduce engineering overhead while simultaneously improving predictive accuracy. The findings suggest that PRAGMA provides a general-purpose representation layer that can be adapted to a wide variety of financial use cases with minimal task-specific training.
Modern financial systems generate vast quantities of transactional and event-level data that encode rich economic signals. This paper presents PRAGMA, a family of foundation models for banking event sequences. Our approach pre-trains a Transformer-based architecture with masked modelling on a large-scale, heterogeneous banking event corpus using a self-supervised objective tailored to the discrete, variable-length nature of financial records. The resulting model supports a wide range of downstream tasks such as credit scoring, fraud detection, and lifetime value prediction: strong performance can be achieved by training a simple linear model on top of the extracted embeddings and can be further improved with lightweight fine-tuning. Through extensive evaluation on downstream tasks, we demonstrate that PRAGMA achieves superior performance across multiple domains directly from raw event sequences, providing a general-purpose representation layer for financial applications.
Alex: Right, and the ablations I'd look for would isolate that. Does removing the periodic embeddings hurt, and does the gain from the log-gap feature hold up? Those are the supporting checks. The main evidence is the cross-domain comparison, and the tokenization is the main reason the authors give for it. [[RP_SECTION:adaptation-and-deployment|Adaptation and deployment]]
Sam: What about distribution shift? Training covers twenty-five months of data, and user behavior drifts.
Alex: The authors treat it as a limitation. They balance how much history they cover against computational constraints, so the training window is a compromise. They do support LoRA fine-tuning, which lets the pretrained model specialize to a task by updating only a small fraction of its parameters. That makes adaptation cheap. It isn't evidence that the base representation stays valid as behavior moves, so temporal robustness remains open.
Sam: There's also an efficiency question. Could this run in something like real-time fraud detection?
Alex: I'd be careful there. The paper describes shard-based batching and sequence packing, which avoid padding overhead. That makes training and processing variable-length histories more efficient. It's a different matter from serving latency for a live fraud decision, and the paper doesn't settle that. [[RP_SECTION:summary-of-findings|Summary of findings]]
Sam: So the reliable parts are the tokenization scheme, the explicit time handling, and consistent gains over task-specific baselines. The universal representation language and the deployment story still need evidence.
Alex: That's how I'd weigh it. The tokenization is the contribution, and the benchmark wins are the evidence for it. Drift and deployment are where a careful referee would push.
Sam: Then the open questions are about time. What happens to the representation after the training window ends?
Alex: If you want the figures and the method choices we skipped, you can generate a deep dive of this paper. The paper has the rest either way.
Sam: Thanks for listening.