Half of Bank AI Projects Never Reach Production

The Southeast Asia Paradox

Southeast Asian banks are spending more on AI than at any point in history, backed by over $30 billion committed to AI-ready data center infrastructure and generally supportive regulators in markets like Singapore and Thailand. Yet roughly half of all AI projects in banking and financial services never make it to production, launched with fanfare as pilots, then quietly shelved. The gap between infrastructure investment and actual returns points to something other than a technology shortfall.

Why This Matters for BFSI and Collections

The article’s central diagnosis is organizational, not technical: AI is still treated as an IT initiative in many institutions, sitting with a technology team or innovation unit and never reaching the frontline staff and customers who would actually use it. One leading Singaporean bank generated $565 million from more than 350 AI use cases in 2024, not by building an impressive demo and waiting for results, but by embedding AI into how daily work actually happens, in credit decisioning and customer engagement across the business. For collections operations specifically, this is a direct warning against treating a voicebot deployment as a technology procurement decision owned by IT, rather than an operational change owned by the collections floor itself. Krungsri’s research reinforces the same point from another angle: AI-driven personalization has been linked to roughly 6% revenue uplift, but only when it is built into everyday infrastructure rather than deployed as a standalone tool sitting apart from how work actually gets done.

What the Numbers Do Not Say Out Loud

The article draws a useful distinction between digital-native banks like GXS and Trust Bank, which can move faster because they carry no legacy technical debt, and traditional banks, which cannot pause operations serving millions of customers to modernize. What it does not spell out explicitly is the trap this creates: faced with legacy complexity, many institutions default to expensive, multi-year replacement programs that accumulate scope creep and stall before delivering any value at all. The more durable path the article points to, extending what already exists rather than displacing it wholesale, is a quieter strategy, and it is precisely the kind of approach that gets less attention than ambitious transformation announcements.

The Practical Read

For collections and BFSI operations evaluating a voicebot deployment, the lesson is not to wait for a full core-system modernization before automating anything. It is to define what success looks like in cost or risk terms before deployment, involve the frontline agents who will actually use the tool in its design, and treat change management as a core part of the rollout rather than an afterthought once the technology is live. A collections voicebot layered onto existing case management and payment systems, with the team that handles hardship calls actually shaping how it behaves, is a more durable bet than a standalone pilot run by IT in isolation.

[Read the full report]

Related Post

Most AI Pilots Do Reach Production

The Finding Metrigy’s AI Organizational Best Practices 2026–27 study, surveying 756 companies globally, pushes back on the common narrative that most AI pilots stall out. Eighty-five percent of companies report fewer than 40% of their pilots fail to reach production, and the largest single group (30.3%) reports a failure rate of just 11 to 20%. […]

When AI Agents Take Action, Mistakes Stop Being Content Errors and Start Being Consequences

The Announcements Google Cloud Summit London 2026 brought a wave of partnership news signaling a shift from AI experimentation to production deployment. HSBC and Google Cloud announced a partnership to deploy more than 200 new AI use cases over the next two years, focused on hyper-personalized wealth management and financial crime detection. Deloitte and Google […]

Enterprises Are Underestimating Multi-Model AI Failure Rates by 2.25x

A new study evaluating 67 frontier models from 21 providers, including GPT-5.5, Claude Opus 4.8, and Gemini 3.1 Pro, found that combining multiple AI models does not create the safety net most enterprises assume. On the MATH-500 benchmark, standard correlation metrics predicted a 2.3% “co-failure rate,” the share of prompts where every model in a […]