AI in banking

How to prioritize agentic AI use cases in banking

27 May 2026
9
mins read

Picking a first agentic AI workflow invites two mistakes. Teams choose the most exciting case, or the easiest one. A scoring method keeps attractiveness and preparedness apart.

Every score, weight, and threshold below is this guide's illustrative method. No one has validated it, so calibrate it to your bank's own risk appetite. Nothing here is legal advice, and nothing here promises a return.

Is the workflow suited to an agent?

Start with the work itself. A fixed rule set handles stable, structured work, so build the deterministic workflow there. Approval thresholds and routing rules are examples.

An agent handles a task that spans several steps and requires reading context to choose among paths. Gathering evidence for a dispute is one example. The nine-workflow overview lists illustrative candidates. The agentic banking explainer defines the term.

Many workflows need both. Backbase places deterministic workflows and agentic work side by side in the Orchestration layer of the AI-native Banking OS. Consequential actions should be attributable and explainable, with human control for judgment and exceptions. See the current Banking OS overview for the wider architecture. The banking operations page shows Customer Operations examples.

Agentic AI in banking scoring method

Score each criterion from 1 to 5 and multiply by its weight. Weights on each scale total 100. Sum the results to get a weighted total between 100 and 500.

Normalize to 0-100 with this formula: (weighted total - 100) / 4. A scale where every criterion scores 3 totals 300 and normalizes to 50.

Score value and readiness on separate scales. Never average them. A high value score says nothing about preparedness.

Label every score with a confidence level. High means measured or tested on real cases. Medium means an owner's estimate from records. Low means assumption or opinion.

Compact rubric

```html
Scale Criterion Weight A score of 5 means
Value Case volume and recurrence 25% High, steady volume
Value Manual effort per case 25% Heavy handling and rework
Value Customer or employee impact 20% Clear pain removed
Value Measurable outcome 20% Baseline and target exist
Value Reuse potential 10% Components serve other workflows
Readiness Data access and quality 25% Complete, current, permissioned
Readiness System connectivity 25% Read and write access tested
Readiness Process clarity 20% Steps and exceptions documented
Readiness Test evidence 20% Tested on real case samples
Readiness Build and run capacity 10% Named, staffed team
```

Six pass-or-fail readiness gates for agentic AI workflows

Gates are pass or fail. Each needs a named person to confirm it. No gate can be overridden by value.

  1. Accountable owner. One named executive owns the outcome and the risk.
  2. Approved boundary. The actions, data, and systems the agent may touch are written down and approved.
  3. Reviewer capacity. Enough trained reviewers can handle escalations and exceptions within service targets.
  4. Lawful purpose and policy review. Your legal, compliance, and policy teams have reviewed the purpose.
  5. Tested suspension and containment. You have tested how to switch the agent off and limit its reach.
  6. Attributable audit trail. Every consequential action links to an actor, evidence, and an approver.

Decision thresholds

Apply these thresholds to the normalized scores and the gate results.

  • Proceed: all six gates pass, value is 60 or higher, readiness is 70 or higher, and no evidence carries Low confidence.
  • Prepare: a credible, named, dated closure plan exists for each failed gate, or readiness is 50-69 with all gates passed.
  • Defer: value is below 60, readiness is below 50, or a failed gate has no credible closure path.

A strong value score never rescues a failed gate. Part 2 sets out the evidence behind each score and works through an example.

Gather evidence before you score

Scores need evidence, and opinion is a weak substitute. Collect three sets of material for each candidate.

  • Common cases. Pull a representative sample of recent work across case types and channels, at normal and peak volumes. Record the manual steps and current error patterns.
  • Edge-case challenge set. Keep this apart from the common cases. Include incomplete documents, conflicting data, unusual amounts, and customers showing signs of vulnerability. Score it separately.
  • System and data checks. List every system the work touches, who owns each, and whether access and data quality hold up.

Common cases show typical performance. The challenge set shows where a person must step in.

Worked example (all scores hypothetical)

Both candidates below are hypothetical. Each score uses (weighted total - 100) / 4. Value raw scores run in the order volume, manual effort, impact, measurable outcome, reuse. Readiness raw scores run data, connectivity, process clarity, test evidence, capacity.

Worked example (all scores hypothetical)
Candidate Value raw scores Value weighted total Value score Readiness raw scores Readiness weighted total Readiness score
A 4, 3, 3, 4, 4 355 63.75 4, 4, 4, 4, 4 400 75
B 3, 5, 5, 3, 3 390 72.5 3, 3, 3, 3, 3 300 50

Candidate A passes every gate under these invented assumptions:

  • A named business owner exists.
  • The approved scope is written down.
  • Reviewers and capacity are adequate.
  • Legal and compliance review is complete.
  • Suspension has been tested.
  • The audit trail is attributable.

On those assumptions, Candidate A is the first candidate to evaluate.

Candidate B scores higher on value but fails two gates: lawful-purpose and policy review, and current reviewer capacity. In this hypothetical, its compliance lead has a dated review scheduled. An approved staffing plan adds two trained reviewers before testing. Each failed gate has a named, dated closure plan, so Prepare applies. A high value score does not clear a gate.

Minimum checklist for controlled production

Passing a score is the start. Production also requires:

  • Gates tested, with results recorded.
  • Evaluation criteria agreed before launch, covering common cases and the challenge set.
  • Monitoring with a named owner.
  • A manual or deterministic fallback.
  • Reviewer capacity for peak volumes.
  • An incident, suspension, containment, and recovery plan.
  • Consequential actions that are attributable and explainable, with people handling judgment and exceptions.

Further reading

Backbase's comparison of agentic AI providers for banks is non-ranked. It draws only on public vendor material and includes no independent product testing.

References

  1. Financial Stability Board. Sound Practices for Responsible Adoption of Artificial Intelligence (AI): Consultation report. June 10, 2026.
  2. NIST. AI Risk Management Framework. AI RMF 1.0.
  3. NIST. AI 600-1: Artificial Intelligence Risk Management Framework, Generative AI Profile. July 2024.
  4. European Banking Authority. Rising application of AI in EU banking and payments sector. September 25, 2025.
  5. European Banking Authority. AI Act implications for the EU banking and payments sector. November 21, 2025.
About the author
Table of contents
Vietnam's AI moment is here
From digital access to the AI "factory"
The missing nervous system: data that can keep up with AI
CLV as the north star metric
Augmented, not automated: keeping humans in the loop