84% of AI Pilots Never Scale: The Problem Is Not the AI

The Finding

Nearly half of enterprises investing heavily in AI are struggling to see meaningful returns, according to Kyndryl’s own Readiness Report, a finding echoed by a separate IBM study showing that just 25% of AI initiatives deliver expected ROI, with only 16% scaling enterprise-wide. In response, Kyndryl has launched Agentic Service Management, a maturity model and implementation framework built specifically to diagnose why AI stalls at the pilot stage rather than reaching production at scale.

Why This Matters for BFSI and Collections

Kyndryl’s framework explicitly names financial services, alongside healthcare and the public sector, as environments where this problem carries the highest stakes. As Kris Lovejoy, the company’s Global Head of Strategy, put it, most enterprise environments were built for people running tickets and tools, not for fleets of autonomous agents executing tasks across hybrid and multi-cloud estates. For BFSI and collections operations, this distinction matters directly: an AI agent handling a payment negotiation or case escalation without clearly defined operational boundaries is not just an efficiency risk; it is a compliance exposure. The framework’s use of ISO 42001, the AI management standard most enterprises have notChoice-wise engaged with yet, signals that governance is becoming a formal benchmark rather than an internal best practice.

What the Numbers Do Not Say Out Loud

The 25% ROI and 16% scaling figures describe an industry-wide pattern, but they obscure a more specific point buried in Kyndryl’s own positioning: the company is applying this framework to its own operations first, running nearly 200 million automations a month across more than 8,000 certified playbooks through its Kyndryl Bridge platform, before offering it to clients. That ordering matters. It suggests the gap between AI pilots and AI outcomes is not primarily a technology maturity problem, since the underlying automation clearly works at scale internally, but an operating-model problem: workflows, approval chains, and oversight structures are still designed for manual, ticket-based work. The uncomfortable implication for BFSI is that buying a more capable AI agent does not close this gap. The bottleneck sits in the surrounding operational architecture, not in the agent itself.

The Practical Read

For collections and BFSI teams evaluating or scaling voice AI, the practical lesson is not to wait for a governance framework before deploying; it is to build the operating model and the deployment together rather than sequentially. A voicebot that can competently handle a payment conversation still needs clearly defined escalation paths, audit trails, and boundaries on what it is authorized to decide versus what requires human sign-off. Retrofitting that structure after a pilot succeeds is exactly the pattern behind the 84% of AI initiatives that never scale. The organizations most likely to move past the pilot stage are the ones treating governance as part of the deployment from day one, not as a compliance step added once the technology is already proven.

[Read the full report]

Related Post

Voice AI Has Crossed the Adoption Threshold: What Changed?

The Shift Underway Enterprise demand for voice AI has moved from cautious experimentation to default deployment in 2026. The voice AI agents market is projected to grow from $2.4 billion in 2024 to $47.5 billion by 2034, representing a 34.8% CAGR. Gartner projects that by year-end 2027, conversational AI will automate roughly 70% of customer […]

Enterprise AI Agents Are Doing More Than Anyone Approved

The Real Question Enterprises Are Asking Most organizations have already decided to deploy AI agents. The open question, according to Joe Davis, EVP of AI Engineering and Delivery at ServiceNow, and Adel El Hallak, VP of Product Management at NVIDIA’s Agentic AI division, is what those agents can actually do, and at most organizations today, […]

McKinsey’s 2026 AI Trust Survey Discovery

The Maturity Gap Behind the Headlines McKinsey’s 2026 AI Trust Maturity Survey, based on responses from roughly 500 organizations collected between December 2025 and January 2026, found the average responsible AI maturity score rose from 2.0 to 2.3 year over year. But that improvement masks an uneven picture: only about 30% of organizations reached a […]