Backbase
Backbase AI Research presents peer-reviewed study, validated across 40+ financial institutions, which finds a model trained to say "I don't know" resolved significantly more customer queries than systems built to always answer, while outperforming GPT-4.1 at up to 50x lower cost.

AMSTERDAM, 30th July, 2026: Backbase, the company behind the AI-native Banking OS, today released the findings of a peer-reviewed study on production-grade banking AI, presented at the 64th Annual Meeting of the Association for Computational Linguistics (ACL 2026). It was led by Denys Katerenchuk, Head of AI Research at Backbase, previously of Google and IBM. It is among the first peer-reviewed accounts of a banking-grade language model measured in live production.
The hardest problem in customer-facing banking AI is what the model does when the evidence isn't there. A system that invents an answer about a fee, a rate, or a policy creates regulatory exposure. But one that declines too often becomes useless. The study shows this trade-off can be engineered out. A 12-billion-parameter model trained to recognize the limits of its own evidence resolved significantly more customer queries in live deployment. It also outperformed GPT-4.1 on the quality and grounding measures that matter most in a regulated environment, at a fraction of the cost.
"Off-the-shelf models tend toward hallucination and sycophancy - confident, agreeable answers even without evidence. That's especially risky in banking, where information is complex, technical, and scattered across dozens of documents," explained Katerenchuk.
The stakes of this disconnect are well documented. McKinsey estimates AI could drive up to 20% in net cost reductions for banks. Yet MIT research found that 95% of enterprise generative AI pilots deliver no measurable P&L impact.
Katerenchuk and his team trained a model to understand the domain and recognize when information is incomplete. They taught it the boundaries of its own knowledge, so it says "I don't know" instead of inventing an answer.
Key findings:
The findings land amid a live industry debate: research published by OpenAI in 2025 found that training and evaluation methods reward confident guessing over admitting uncertainty. The Backbase study offers production evidence of the alternative: a model rewarded for honesty, measured against real customers.
"For three years, the AI industry has rewarded models for speed and confidence. Banking has rewarded itself for the same thing for three decades. Saying 'I don't know' got treated as a weakness, not a feature," said Jouk Pleiter, Founder and CEO of Backbase. "Our research shows the opposite: a model that knows the limits of its own evidence earns more trust, not less."
"2026 is the year agentic workflows go live in regulated environments, but none of it works unless the model knows what it doesn't know," added Pleiter.
This study is also the first published work from Backbase AI Research - the team that joined Backbase through its acquisition of Kasisto. The group focuses on the specific problems of AI in banking, publishing peer-reviewed research openly and moving findings directly into production.
- ENDS -
Notes:

Backbase built the AI-native Banking OS - the operating system that turns fragmented banking operations into a Unified Frontline. Customers, employees, and AI agents work as one across digital channels, front-office, and operations.
Backbase was founded in 2003 by Jouk Pleiter and is headquartered in Amsterdam, with teams across North America, Europe, the Middle East, Asia-Pacific, Africa and Latin America. 120+ leading banks run on Backbase across Retail, SMB & Commercial, Private Banking, and Wealth Management.