30 July 2026

Backbase

Banking AI trained to admit uncertainty resolves more customer queries at a fraction of the cost

Backbase AI Research presents peer-reviewed study, validated across 40+ financial institutions, which finds a model trained to say "I don't know" resolved significantly more customer queries than systems built to always answer, while outperforming GPT-4.1 at up to 50x lower cost.

AMSTERDAM, 30th July, 2026: Backbase, the company behind the AI-native Banking OS, today released the findings of a peer-reviewed study on production-grade banking AI, presented at the 64th Annual Meeting of the Association for Computational Linguistics (ACL 2026). It was led by Denys Katerenchuk, Head of AI Research at Backbase, previously of Google and IBM. It is among the first peer-reviewed accounts of a banking-grade language model measured in live production.

The hardest problem in customer-facing banking AI is what the model does when the evidence isn't there. A system that invents an answer about a fee, a rate, or a policy creates regulatory exposure. But one that declines too often becomes useless. The study shows this trade-off can be engineered out. A 12-billion-parameter model trained to recognize the limits of its own evidence resolved significantly more customer queries in live deployment. It also outperformed GPT-4.1 on the quality and grounding measures that matter most in a regulated environment, at a fraction of the cost.

"Off-the-shelf models tend toward hallucination and sycophancy - confident, agreeable answers even without evidence. That's especially risky in banking, where information is complex, technical, and scattered across dozens of documents," explained Katerenchuk.

The stakes of this disconnect are well documented. McKinsey estimates AI could drive up to 20% in net cost reductions for banks. Yet MIT research found that 95% of enterprise generative AI pilots deliver no measurable P&L impact.

Katerenchuk and his team trained a model to understand the domain and recognize when information is incomplete. They taught it the boundaries of its own knowledge, so it says "I don't know" instead of inventing an answer.

Key findings:

  • Honesty can be engineered: The model was trained on a dataset in which 22% of examples had no correct answer. That taught it the right response was an explicit "I don't know." It reached a 12% refusal rate, higher than the untuned base model's 4.3% and lower than GPT-4.1's 20.2%. The base model answered confidently even without evidence. GPT-4.1 declined questions it could have safely answered.
  • The honest model solved more customer problems: Over seven months at a large US financial institution, query resolution rose 7.1 percentage points across 3,297 sampled queries. That's a statistically significant gain. It came despite the model refusing nearly three times as often as its base.
  • A model a fraction of the size beat GPT-4.1: The model scored higher on independent evaluation: 6.21 versus 5.72 for GPT-4.1, on a 10-point scale. Citation grounding improved by 2.3 points, with answers cited directly to source documents. It also produced stronger results on FinanceBench, a public benchmark of SEC filing questions.
  • The economics undercut frontier models by an order of magnitude: Roughly $0.001 per query on a single GPU, 20-50x cheaper than GPT-4.1 and 3-5x faster. Training cost around $1,800.
  • Data order mattered more than data volume: On identical data, teaching general financial language first and calibrated refusal second produced the best model. Pooling everything at once collapsed answer quality by more than 40% and pushed refusals to nearly half of all queries.

The findings land amid a live industry debate: research published by OpenAI in 2025 found that training and evaluation methods reward confident guessing over admitting uncertainty. The Backbase study offers production evidence of the alternative: a model rewarded for honesty, measured against real customers.

"For three years, the AI industry has rewarded models for speed and confidence. Banking has rewarded itself for the same thing for three decades. Saying 'I don't know' got treated as a weakness, not a feature," said Jouk Pleiter, Founder and CEO of Backbase. "Our research shows the opposite: a model that knows the limits of its own evidence earns more trust, not less."

"2026 is the year agentic workflows go live in regulated environments, but none of it works unless the model knows what it doesn't know," added Pleiter.

This study is also the first published work from Backbase AI Research - the team that joined Backbase through its acquisition of Kasisto. The group focuses on the specific problems of AI in banking, publishing peer-reviewed research openly and moving findings directly into production.

- ENDS -

Notes:

About the author
Backbase
Backbase is on a mission to to put bankers back in the driver’s seat.

Backbase built the AI-native Banking OS - the operating system that turns fragmented banking operations into a Unified Frontline. Customers, employees, and AI agents work as one across digital channels, front-office, and operations.

Backbase was founded in 2003 by Jouk Pleiter and is headquartered in Amsterdam, with teams across North America, Europe, the Middle East, Asia-Pacific, Africa and Latin America. 120+ leading banks run on Backbase across Retail, SMB & Commercial, Private Banking, and Wealth Management.

Table of contents
Example H2
Example H3
Example H4
Example H5
Example H6