Agentic AI is moving from pilots into live banking work. The European Banking Authority reports that 55% of surveyed EU banks already use general-purpose or agentic AI in consumer-facing processes.
That raises a buying question a model demo cannot answer: what happens after the agent decides? Architecture answers it. It determines what the agent can see, what it may do, how the work runs, and what gets recorded.
This article proposes a buyer-review model for that conversation. We suggest buyers review four criteria for every consequential action. A consequential action moves money, changes a customer record, or commits the bank to a decision.
- Bounded: the action sits inside authority the bank defines.
- Attributable and explainable: the bank can show which human, agent, or service acted, and why the action was allowed.
- Recoverable: the bank knows what state remains if a step fails.
- Under human control: people keep control of judgment and exceptions, and a named function owns the process outcome.
These review criteria are Backbase's proposed model for structuring your own evaluation.
Why does architecture decide whether an agent can act in a bank?
A model can respond to a prompt and propose an answer. Running a bank asks for more. The bank needs memory of what happened, knowledge of who may act, coordination across systems, and a record of why each consequential action was allowed.
Several voices point the same way. Deloitte's UK financial services team argues that some banks build agents first and retrofit controls later. It recommends designing the operating model first, then deriving build requirements from it. EY makes a related argument: banks will need an orchestration layer that enforces guardrails and provides auditability.
Backbase Banking OS runs an Authority Layer. It decides whether any human or machine actor may perform a consequential action. Every action is evaluated against identity, policy, evidence, and limits. Authority contracts define autonomy level, monetary limit, confidence threshold, required evidence, and human-approval triggers, per action and actor. Each authorized action yields a Decision Token, a real-time authorization carrying scope, constraints, and reason codes. The platform records what happened and why, with human control for judgment and exceptions. The bank sets autonomy by domain, use case, actor, and action. Tightening, suspending, or revoking autonomy leaves the underlying model unchanged.
What is the buyer-review model?
The proposed model has four lenses: Context, Authority, Execution, and Evidence. Apply them to one named use case. A dispute, a payment inquiry, or a KYC remediation case works well.
Context: what does the agent know, and where does that knowledge come from?
- Which systems feed the agent, and how current is each source?
- Does state persist across handoffs between customer, employee, and agent?
- What happens when two sources disagree?
- Does memory extend beyond the current prompt?
Authority: who defines what the agent may do?
- Can the bank set limits per actor and per action?
- Which conditions force human approval, such as an amount over the limit, conflicting evidence, a fraud signal, or a vulnerable customer?
- Can the bank tighten, suspend, or revoke autonomy without changing the model?
Recommendation: ask vendors to describe autonomy as a ladder the bank controls. One way to frame it runs from observe, to recommend, to prepare, to act with approval, to act within limits. The bank sets the level for each domain, use case, actor, and action.
Execution: how does the work run across systems?
- Which steps run as deterministic workflow, and which use an agent's judgment?
- What state is preserved if a step fails midway?
- Can the bank retry, roll back, or route the case to an employee?
- Do systems of record remain the source of truth?
Evidence: what record exists after the action?
- Can a risk or audit team reconstruct who or what acted, on what authority, with what evidence, and under which policy version?
- Is the record produced as the work runs, or assembled afterward?
- Can the bank explain the outcome to the customer in plain language?
Recommendation: weight Authority and Evidence more heavily as the agent's autonomy rises. A read-only assistant needs lighter review than an agent that can move money.
Which architecture patterns should a bank compare?
The four patterns below are examples to compare, and a given product may combine several. Vendors may claim lower integration effort for any of them. Treat each claim as a question to verify in a proof of value.
Pattern one: an agent attached to a single channel or application.
- Can it reach the other systems the process touches?
- Does context follow the customer into another channel or to an employee?
- Who owns the action log?
Pattern two: agent steps inside an existing workflow or case tool.
- Can the workflow engine enforce authority limits at the agent step?
- How are exceptions routed, and to whom?
- If a process follows fixed rules, would a deterministic workflow do the same job?
Pattern three: an agent built on a framework by the bank's own engineering team.
- Who builds and maintains identity, authority, evaluation, and logging components?
- What does upkeep cost over time?
- How easily can the team change models without reworking controls?
Pattern four: a governed layer above systems of record that coordinates people, agents, and software.
- How does it synchronize with cores, CRMs, and payment systems?
- Does it add another system to secure and operate?
- Can the risk team inspect its decisions directly?
If your target process mixes fixed rules and judgment, you may need more than one pattern. If most of a process follows fixed rules, a deterministic workflow may be enough. The deterministic vs agentic comparison covers that decision.
What evidence should a buyer request?
Ask every vendor, and your own engineering team, for artifacts such as configuration files, sample records, and test results. The matrix below turns the four criteria into requests. The NIST mapping is our interpretation.
Recommendation: assign each row to a named reviewer in risk, operations, or technology before the vendor session. Rows without an owner tend to go unanswered.
How do NIST and the FSB fit into this review?
Two public documents give buyers a common vocabulary for the model above.
NIST describes its AI Risk Management Framework as intended for voluntary use. AI RMF 1.0 organizes AI risk management into four functions: Govern, Map, Measure, and Manage. NIST positions Govern as cross-cutting. NIST also notes that AI RMF 1.0 is being revised, so check the current version.
The Financial Stability Board published a consultation report on June 10, 2026. It proposes a menu of 12 sound practices for organization-wide AI governance and lifecycle management. It is a consultation, so its 12 practices are proposals. Use both as reference points for your own control set.
What does a worked example look like?
This walkthrough is hypothetical. Its figures are invented for illustration.
A customer reports a duplicate card charge of €86.40 in the mobile app. The bank has configured the dispute process as follows.
- Intent: the customer describes the problem in conversation, and the system identifies a dispute.
- Context: the system retrieves the card transactions, any prior disputes, and the case state. An employee who picks up the case later sees the same picture.
- Proposal: the agent proposes refunding the duplicate charge.
- Authority check: the bank's rule allows automatic action up to €250 when confidence is at least 0.92, the evidence is consistent, and no fraud or vulnerability flag is present.
- Execution: the card service issues the refund, the case updates, and the customer is told.
- Record: the system stores the actor, the authority applied, the evidence used, and the reason codes.
Three variations show where the review earns its keep:
- Higher amount: a €900 duplicate charge exceeds the limit, so the case routes to an employee for approval.
- Conflicting evidence: the merchant record contradicts the customer's account, so the case routes to a person with the full timeline.
- System failure: the card service times out mid-case, so the case state is preserved and the system retries or routes to operations without issuing a second refund.
Ask each vendor to run this scenario on your use case. Then ask to see the records it produces.
How does a phased maturity model fit the review?
Backbase's whitepaper, Agentic Banking: A Phased Blueprint to Elastic Operations, presents Crawl, Walk, Run as capability-maturity phases. They run across Conversational Banking, Customer Operations, and proactive advisory, which Backbase now calls Relationship Intelligence.
The phases describe how capability matures. The landing page also covers outcome measurement, so a reviewer can ask which business results each phase should move. Use it as a supporting resource. A bank can borrow the idea to scale review depth with autonomy: the more an agent can do on its own, the deeper the review should go.
How can a bank test production readiness?
Run these tests before an agent acts on live customer work. Only an executed test produces a result, so keep the evidence from each run.
For the measurement test, we suggest tracking business outcomes. Control measures such as policy adherence, human exceptions, and evidence completeness belong beside service and growth metrics.
Where does Backbase stand on this?
Backbase's position is that AI can reason, and Banking OS makes it safe to act. The AI-native Banking OS provides Context, Authority, and Execution. The model contributes intelligence, and Banking OS owns the outcome. It is the Control Plane of the Unified Frontline, where customers, employees, and AI agents work from the same context, policies, and execution layer. Digital Banking creates the interface, Agentic Banking resolves the work, and Banking OS powers both.
Within Agentic Banking, Conversational Banking and Relationship Intelligence engage and guide customers, and Customer Operations resolves the work. The Banking OS sits above systems of record, which remain intact. The bank sets autonomy levels, and consequential actions stay attributable and explainable, with human control for judgment and exceptions.
FAQs
Is the buyer-review model an industry standard?
No. It is a proposed model for structuring your own evaluation. Adapt the lenses and evidence requests to your risk appetite, your regulators, and the use case under review.
Is the NIST AI RMF mandatory for banks?
NIST describes AI RMF 1.0 as intended for voluntary use. Its functions are Govern, Map, Measure, and Manage. Your regulators or internal policy may still require specific practices, so confirm with compliance.
Is the FSB report binding regulation?
No. The June 10, 2026 document is a consultation report that proposes sound practices for organization-wide AI governance and lifecycle management. It signals supervisory thinking, and its content may change.
Does an agent need a named person approving every action?
Don't make named-person approval of every action a blanket rule. Consequential actions should be attributable and explainable. Human control should cover judgment and exceptions. Each process should have an accountable owner function. The bank sets the autonomy level per use case and action.
When is a deterministic workflow enough?
If the target process follows fixed rules, a deterministic workflow may do the job with less review burden. Agents add most where judgment is needed. See the deterministic vs agentic comparison for the technical detail.
Where should a bank start?
Pick one consequential use case with a clear owner and measurable result. The use-case prioritization guide covers selection. The agentic workflows guide covers how the work runs.
How do I compare vendors against this model?
Send each vendor the same walkthrough and the same evidence matrix, and score the artifacts they return. The provider comparison covers vendor evaluation in more depth.
Does a bank need to replace its core to use this?
Not by default. Whether your architecture requires it depends on the pattern you choose. Ask each vendor how it works with cores, CRMs, and payment systems that stay in place.
For a supporting view of capability maturity and outcome measurement, download the whitepaper, Agentic Banking: A Phased Blueprint to Elastic Operations.
