The Agentic Bank Model: Understanding PRAGMA
And what it tells us about the next decade of AI in banking
In April, thirteen Revolut engineers released PRAGMA, a family of foundation models trained on 24 billion banking events from 26 million customers across 111 countries. It produces no text, answers no prompts, holds no conversation.
It reads a person’s financial history and returns a single dense vector, a compressed mathematical portrait of who that customer is.
That quiet capability is the foundation the next decade of banking agents will be built on.
This week’s piece wants to unpack what PRAGMA is, how it works, what it can and can’t do. Then, the part that matters most for anyone running a bank or a payments business: what changes when an agent can read one of these models and act on what it sees.
What PRAGMA is
PRAGMA is an encoder. You feed it a sequence of a customer’s transactions, card events, and account activity, and it outputs an embedding: a list of numbers that places that customer in a space where financially similar people sit close together.
The model comes as a family, from a 10-million-parameter version to a billion-parameter one. The same backbone supports, in the paper’s own words, “credit scoring, fraud detection, and lifetime value prediction”, among the six downstream tasks it evaluates. One model, trained once, feeding many decisions. That is the entire idea.
How it works: teaching a model to read money
The clever part of PRAGMA is how it represents an amount. A general model splits a figure into its digits, which throws away magnitude and ordering: the model sees the characters in “$1,000” and learns almost nothing about how big a thousand euros actually is. Revolut uses percentile tokenization instead.
Each amount is mapped to a percentile bucket, drawn from the distribution of transaction sizes across Revolut’s whole dataset, so a $1,000 transaction sits semantically close to a $900 one and far from a $5 coffee.
This the model’s most important innovation. It lets the model read an amount in context, the way a good underwriter does intuitively. What the model captures is the shape of a financial life.
How they built it
What should catch your eye is the cost of entry. It’s not high at all.
The smallest version of PRAGMA finished training in about two days on 16 GPUs, after learning from 25 months of customer behavior. Once the model exists, pointing it at a new job (churn this quarter, fraud the next) means updating only 2 to 4 percent of its weights, not rebuilding from scratch. Revolut says that cuts the time to ship a new model by 3 to 5 times.
Pavel Nesterov, one of the paper’s authors, framed the experiment plainly to Simon Taylor: “can we beat the entire data science team by taking a pre-trained model and pressing a button?” The lifts say largely yes: 130% better accuracy at spotting rare cases of credit-default risk, a 64.7% improvement in fraud recall, and a 40.5% gain in product-recommendation relevance, all from one shared backbone.1

Whether your institution can assemble the data, the infrastructure, and the talent to ship one is the open question, and it sits in operations, not research.
From prediction to action: the agentic substrate
Here is the mental model I use when a client asks where this is heading. Picture three layers stacked on top of each other.
Layer one: the substrate
At the bottom sits the embedding model itself, the perception organ of the system. PRAGMA, Nubank’s nuFormer, Visa’s TREASURE all live here. Their only job is to turn raw behavior into a representation: a precise, current read on who a customer is, financially, in a form that downstream systems can consume.
Layer two: the prediction layer
Freeze the backbone and attach a lightweight head, and you get a prediction. Will this customer default? Is this transaction fraud? What is their lifetime value, and are they about to leave? This is where every bank already operates, and where PRAGMA’s published results sit. It’s familiar territory, made cheaper and sharper by a shared foundation underneath. Most institutions will stop here and still capture real value.
Layer three: the action layer
The third layer is where the customer-experience step function actually happens, and it’s the one few companies have built yet.
I’ve argued before that an AI agent is defined by four things: the environment it operates in, how it perceives that environment, its capacity to reason about what it perceives, and the actions it can take.
A behavioral foundation model is a near-perfect perception layer for a banking agent. It hands the agent a precise, real-time understanding of the customer. Reasoning sits in the orchestration model on top. And action is where the value gets created.
What does action mean concretely? An agent sitting on this stack could:
Execute a domestic or international transfer between the customer’s own accounts or to a third party.
Move part of an available credit-card limit into a current account, or back the other way.
Raise a credit limit, within pre-approved bounds.
Tailor a loan offer inside a single conversation and, on acceptance, formalize it in the core system.
Contextually offer and contract a new product, a card or a brokerage account, the instant a customer signals intent.
Act to retain a customer the model has flagged as likely to churn.
The list runs into the hundreds. None of it works without two things.
First, the institution has to expose the right internal endpoints, securely and at scale, so an agent can actually move money and change limits rather than only talk about them.
Second, the agent’s developers have to equip it with the right tools: an instant-SEPA-transfer tool, a tailor-loan-conditions tool, a contract-product tool, each one a thin, audited wrapper that calls those endpoints to finalize an action in full or in part. The model perceives. The tools act.
Which is why the infrastructure and data-availability gap is the key. The institutions that close the endpoint-and-tooling gap fastest are the ones that turn these models from clever predictions into agents that do things.
The limitation every executive should know
Before this reaches a steering committee, name the limit, because it’s real.
On anti-money-laundering, PRAGMA underperforms Revolut’s own production baseline by 47.1%. The reason is structural: the model reads each customer in isolation, and money laundering lives in the relationships between accounts.
It can’t see the chain. Anything relational, laundering rings, mule networks, coordinated fraud, needs a graph, not a sequence.
A behavioral foundation model like this wins as a shared perception layer feeding specialized heads and agents, not as a single oracle making every call.
The PRAGMA paper concedes as much in its own architecture: it attaches linear and LoRA (Low-Rank Adaptation) heads on top of the core model precisely because the embedding alone doesn’t cover it all.
Treat the substrate as the customer’s representation, and keep specialists and graphs for the hardest relational and regulated decisions.
Architected that way, the limitation is an argument for layering, not a reason to wait.
What it means for you
The substrate is the same customer vector everywhere. What changes by business line is the head you attach and the action you let an agent take.
For issuers
Issuing is where this bites first, because card and account activity is exactly the event stream these models read best. The 64.7% fraud-recall lift lands here directly. So does the agentic upside: an agent that perceives a customer through the embedding can raise a pre-approved limit on request, shift part of a credit line into a current account, handle a stolen-card report end to end, and offer a card upgrade the moment usage patterns justify it, contracting it inside the same conversation.
For consumer finance
Consumer finance is where both the upside and the regulatory weight concentrate. The 2.3 times improvement in default-risk accuracy is a lending number, and an agent could tailor a loan offer in one conversation and formalize it on acceptance.
The catch is regulatory. Under the EU AI Act, credit scoring sits in the high-risk tier, with obligations landing 2 August 2026, while fraud detection is exempted. Explainability of a billion-parameter embedding is an unsolved problem, and GDPR still bars solely-automated decisions that significantly affect someone.
The pragmatic architecture keeps a human in the loop and an auditable specialist head on the decision, while the model does the perceiving.
For acquirers
The same pattern already shipped on the acceptance side: Stripe’s payments foundation model lifted card-testing detection from 59% to 97% with no increase in false positives, across more than a trillion dollars of annual volume.
For broader banking use cases, investments, insurance, mortgages, business banking, the recipe holds: one customer vector, different heads, different actions. A model that predicts churn can feed an agent that rebalances a portfolio, flags an underinsured life event, or pre-qualifies a mortgage top-up. The architecture stays constant while the use cases multiply.
What you could do
The question for any bank or payments company stops being whether you’ll need a behavioral foundation model. It becomes how you’ll get one: build it, partner for it, or buy it.
Nubank built nuFormer, a 330-million-parameter transformer serving its ~131 million customers. Stripe shipped their own on the acceptance side. PayPal went a step further and wrapped its model in a multi-agent system where separate agents handle reasoning, checkout, and fraud (exactly the action layer in production). T
Visa, also built its own. TREASURE, our payment foundation model, lifts abnormal-behavior detection by 111% over production systems as a standalone model and improves recommendation models by 104% as an embedding provider.
I’ll write about it properly in the coming weeks.
For now, decide which layer you intend to compete on. If you want help building a behavioral foundation model, wiring it to a prediction layer, and exposing the secure endpoints and tools that turn it into an agent that actually acts for your customers, that’s the work we do in the AI Labs at Visa Consulting. Get in touch.
Caveat: Revolut publishes only relative lifts, not absolute baselines, which it calls commercially sensitive. Read the magnitudes as directional.



