All resources
Data & AI · 7 min

Why payments needs vertical AI, not a generic LLM

Why a model trained on decline codes and scheme rules beats a general-purpose LLM in the authorization path.

Every payments company now has an AI story. Most of them are a general-purpose language model placed next to a dashboard, asked to summarise what already happened. That is a useful product. It is not the same thing as a model that participates in the authorization decision, and conflating the two hides the constraint that actually decides the architecture.

The latency budget settles the argument

An authorization has a total budget of roughly three seconds before the cardholder feels it and the checkout starts shedding conversions. Inside that budget sit the merchant's own stack, the gateway, the acquirer, the scheme network, and the issuer's decision. Anything that inserts itself into that path gets tens of milliseconds, not hundreds.

A general-purpose LLM answers in several hundred milliseconds at best, seconds under load, with a variance profile that is fine for a chat interface and disqualifying in the auth path. This is not a matter of a faster GPU. It is the wrong shape of computation for the deadline.

The knowledge is structured, not textual

What the decision needs is not prose. It is a set of relationships that live in tables and histories:

  • What 05 means from this issuer at this amount — a code whose meaning is learned per-issuer, not read from a spec.
  • Which acquirer has historically been approved by this BIN range for this MCC.
  • Whether an SCA exemption is available for this amount, country, and cumulative spend — rules that differ by market and change on regulatory timelines.
  • Whether a network token performs better than the underlying PAN for this issuer, which is frequently true and occasionally reversed.
  • How this issuer responds to 3DS DataOnly versus a full challenge.

None of that is well served by next-token prediction over text. It is served by a model over structured features, trained on outcomes, that can be evaluated against a holdout and retrained when issuer behaviour drifts.

Domain knowledge is the moat, and it expires

Scheme rules are not static. Interchange tables get revised, SCA guidance changes by market, token mandates arrive with deadlines, and issuers quietly change risk thresholds. A model trained on payments has to be maintained against a moving specification.

That maintenance burden is why a generic model does not close the gap by getting larger. The gap is not reasoning capacity. It is knowing that a particular issuer in a particular market started declining a particular amount band last month — a fact that exists only in recent outcome data.

What "vertical" should mean concretely

The word gets used loosely. A defensible version means at least: features derived from payment-specific entities rather than generic tabular columns; a training objective that is the actual business outcome (approval, cost, recovered revenue) rather than a proxy; and evaluation on cohorts that reflect real mix — by issuer country, card product, MCC, and token type — because an aggregate accuracy number hides the segments where the model is worst.

If a vendor cannot tell you which segments their model underperforms on, they have not measured it that way.

See these patterns in your own traffic

Apex analyzes every transaction against the decline, routing, and cost signals described here.

Request a demo