AI-native, or not at all
AI earns its money by removing the handoffs between teams. Bolt it onto one step of an unchanged chain, and the organisation ends up blaming the model for a failure of design.
- Most enterprise AI never pays back, and the big surveys blame people and process far more than models.
- The firms that do get paid redesign whole workflows. It was the strongest predictor of EBIT impact in McKinsey's survey.
- Most of a process is waiting between handoffs. Speeding up one step saves minutes. Removing handoffs saves days.
Most enterprise AI programmes follow the same arc. A team picks one step in a long process, puts a model on it, and proves in a pilot that the step now runs faster. A year later the P&L looks the same, the pilot is quietly shelved, and the verdict in the boardroom is that AI doesn’t work.
The verdict is wrong, and the evidence for why is now thick. The model was rarely the problem. The half-measure was.
The failure is real, and it is not technical
MIT’s NANDA initiative reviewed more than 300 public AI deployments, interviewed 52 organisations and surveyed 153 senior leaders in the first half of 2025. Its headline: against $30–40 billion of enterprise spending, 95% of organisations were getting zero measurable return.
Of enterprise-grade tools, custom or vendor-sold, 60% of firms evaluated one, 20% piloted, and 5% reached production. The authors are explicit that the barrier “is not infrastructure, regulation, or talent. It is learning.” Most tools never adapt to how the work around them actually runs.1
Other surveys land in the same place from different directions. S&P Global found that 42% of companies said they were scrapping most of their AI initiatives, up from 17% a year earlier, with the average firm scrapping 46% of proofs of concept before production.5 BCG’s survey of 1,000 executives across 59 countries found 74% had yet to show tangible value.4 McKinsey’s 2024 survey found over 80% of respondents saw no tangible enterprise-level EBIT impact from generative AI.2
None of these reports blames the models. BCG puts numbers on it: roughly 70% of the challenge is people and process, 20% is technology, and 10% is the algorithm.4
What separates the firms that do get paid
McKinsey tested 25 organisational attributes against reported EBIT impact from generative AI. The one with the largest effect was the redesign of workflows. Only 21% of organisations using gen AI had fundamentally redesigned even some of them.2
The gap has widened since. In McKinsey’s latest survey, the 6% of respondents it classes as AI high performers (5% or more of EBIT attributable to AI) are nearly three times as likely to say they have fundamentally redesigned individual workflows.3
Adoption is near-universal. Redesign is what is scarce, and redesign is what correlates with money.
In many cases, rethinking workflows with agentic AI from the ground up is the ideal path to successful implementation.
Gartner, predicting that over 40% of agentic AI projects will be cancelled by end-20276
Why the handoff is where the money is
In most knowledge processes, the work itself is a small fraction of the elapsed time. The rest is waiting: for the next team’s queue, for a re-keyed form, for someone to rebuild context that the previous step already had.
Michael Hammer documented this in 1990. At Mutual Benefit Life, an insurance application passed through 30 steps, five departments and 19 people, taking 5 to 25 days. One insurer estimated that an application spent 22 days in process and received 17 minutes of actual work.7 Try the three scenarios below.
Where the time goes in a handoff process
One insurer’s application: 22 days in process, 17 minutes of actual work. Everything else is waiting between handoffs.
Elapsed and work times from Hammer (1990)7. The two-minute saving is illustrative.
Speeding up any single step in that chain by half would have saved minutes out of weeks. Collapsing the chain into one case manager cut turnaround to as little as four hours and doubled productivity.7
Ford’s accounts payable department tells the same story. Ford planned to automate the existing process and cut its 500-person headcount by 20%. Then it saw that Mazda ran the function with five people. Ford removed the invoice-matching handoff altogether and cut headcount by 75%.7 Hammer’s title was the lesson: don’t automate, obliterate.
A step-level AI tool attacks the 17 minutes. An AI-native process attacks the 22 days.
What bolting on looks like in the data
Software delivery. Google’s 2024 DORA report found that about three-quarters of developers use AI and say it makes them more productive. Yet for every 25% increase in AI adoption, delivery throughput fell an estimated 1.5% and delivery stability fell 7.2%.8 Code got written faster and then piled up at review, test and release, which had not changed.
The Danish labour market. Humlum and Vestergaard linked chatbot adoption surveys to administrative records for exposed occupations. Adoption was widespread and workers reported gains, but two years after ChatGPT’s launch the effect on earnings and recorded hours was a precise null, ruling out effects above 2%.9 Individual tasks sped up; the organisations around them did not convert that into output.
Klarna. In February 2024 Klarna reported that its AI assistant handled two-thirds of customer chats in its first month, did the work of 700 full-time agents, cut resolution time from 11 minutes to under 2, and would add an estimated $40 million to profit.10 By May 2025 its CEO told Bloomberg the company had “gone too far” on cost and was hiring human agents again.11 One step was swapped out; the escalation path, the product issues generating the contacts and the service design around it were not.
What AI-native looks like
Ant Group’s MYbank was built with no loan officers in the chain. Its “3-1-0” model means three minutes to apply, one second to approve, zero human interaction. By 2020 it had lent to more than 20 million small businesses at a default rate of about 1%.12
No incumbent bank reaches those economics by adding a credit-scoring model to a branch-based process, because the cost sits in the handoffs between origination, underwriting, approval and disbursement, and MYbank has none.
| Bolted on | AI-native | |
|---|---|---|
| Unit of change | One task in an existing chain | The end-to-end outcome |
| What gets faster | Touch time (minutes) | Elapsed time (days) |
| Handoffs | Unchanged, plus a new one to the model | Removed |
| Evidence | DORA: throughput −1.5%, stability −7.2%8 | MYbank: 1-second approval, ~1% default12 |
| Typical verdict | “AI doesn’t work” | New cost structure |
We have run this experiment before
Paul David’s study of factory electrification is the closest historical parallel. Central power stations opened in the early 1880s, yet by 1899 electric motors supplied under 5% of factory mechanical drive, and productivity did not move until the 1920s.
Early adopters did the equivalent of a bolt-on: they replaced the steam engine with electric motors driving the same shafts and belts, in the same multi-storey layout. The gains arrived only when factories were rebuilt around a motor on each machine. That redesign accounts for roughly half of the five-point acceleration in US manufacturing productivity growth in the 1920s.13
Brynjolfsson, Rock and Syverson generalise this as the productivity J-curve: general-purpose technologies “enable and require significant complementary investments, including co-invention of new processes, products, business models and human capital.” Until those are made, measured productivity dips.14 A firm that buys the technology and declines the redesign pays for the dip and never reaches the upswing.
How the wrong conclusion gets drawn
The sequence is predictable. A step-level pilot produces a real local gain. Rolling it out needs integration, retraining, new controls and a new exception path, which are fixed costs. The local gain is capped by everything upstream and downstream that did not change.
The business case fails, and the post-mortem names the most visible new thing in the room, which is the model. The 42% abandonment rate5 is largely this loop running at scale.
What follows
- Pick outcomes, not steps. Scope by “order to cash” or “claim to settlement”, never by “summarise the document”.
- Measure elapsed time and handoff count before measuring model accuracy.
- Give one owner the whole chain. If nobody can remove a handoff, nobody can capture the value.
- Fund fewer things fully. One redesigned process beats ten assisted steps.
Organisations will keep concluding that AI doesn’t work for as long as they keep testing it in a form that cannot work. The choice is the one factory owners faced a century ago: rebuild the floor around the motor, or keep the belts and wonder where the productivity went.
References
- MIT NANDA, The GenAI Divide: State of AI in Business 2025 (2025). PDF
- McKinsey, The state of AI: How organizations are rewiring to capture value (March 2025; survey July 2024, n=1,491). Link
- McKinsey, The state of AI, latest global survey. Link
- BCG, Where’s the Value in AI? (October 2024; n=1,000, 59 countries). Press release
- S&P Global Market Intelligence survey, reported in CIO Dive, “AI project failure rates are on the rise” (March 2025). Link
- Gartner, “Over 40% of Agentic AI Projects Will Be Canceled by End of 2027” (25 June 2025). Link
- M. Hammer, “Reengineering Work: Don’t Automate, Obliterate,” Harvard Business Review, July–August 1990. PDF
- Google Cloud DORA, Accelerate State of DevOps Report 2024; figures as summarised by RedMonk. Link
- A. Humlum and E. Vestergaard, “Large Language Models, Small Labor Market Effects,” NBER Working Paper 33777 (2025). Link
- Klarna, “Klarna AI assistant handles two-thirds of customer service chats in its first month” (27 February 2024). Link
- CX Dive, “Klarna changes its AI tune and again recruits humans for customer service” (May 2025), citing Bloomberg. Link
- IFC, MYbank case note (August 2020). PDF
- P. A. David, “The Dynamo and the Computer: An Historical Perspective on the Modern Productivity Paradox,” American Economic Review 80(2), 1990, pp. 355–361. PDF
- E. Brynjolfsson, D. Rock and C. Syverson, “The Productivity J-Curve,” American Economic Journal: Macroeconomics 13(1), 2021, pp. 333–372. Link
- E. Brynjolfsson, D. Li and L. Raymond, “Generative AI at Work,” NBER Working Paper 31161 (2023). NBER Digest