Anand Abhishek.
← Essay · Enterprise AI

There is no half-hearted AI

Data, pipelines, models and adoption multiply each other. Four layers that are each good enough compound into a result nobody trusts.

The short version
  • Data, pipelines, models and adoption multiply. A weak layer discounts every strong one.
  • A demo tests the model with the other three layers set to perfect. That is why demos mislead.
  • Fund the weakest layer first, and judge pilots end to end with real data, real pipelines and real users.

Most AI programmes are budgeted as if the layers add up. Spend a bit on data, a bit on engineering, most of it on the model, and something on rollout, and the total should be roughly the sum of the parts. A weak layer costs you its own share and no more.

That is the wrong arithmetic. The layers multiply. The output of each is the input of the next, so a weakness anywhere discounts everything else. A brilliant model on unreliable data is an unreliable system. A reliable system nobody uses returns nothing.

Set the quality of each layer

If the layers added up (average)
90%
Because they multiply
66%

Every layer is at 90%. Notice how far below 90% the outcome sits. Now lower any one layer and watch what happens.

Illustrative arithmetic after Kremer's O-ring model1. Real layers are not literal probabilities, but the direction holds.

The economics of a chain

Michael Kremer formalised this in 1993 and named it after the component that destroyed the space shuttle Challenger. In his O-ring model, production consists of several tasks, a mistake in any one of them can ruin the whole, and so output is the product of the quality at each step, not the sum.1

Two things follow. A single weak link caps the value of every strong one. And it pays to match quality across steps: excellence in one place is wasted next to mediocrity in another.

An AI system is an O-ring process. The evidence, layer by layer:

3%of companies’ data quality scores met a basic “acceptable” standard2
92%of AI practitioners had experienced a compounding data failure downstream3
80%+of AI projects fail, twice the rate of other IT projects (RAND)6

Layer one: data

Nagle, Redman and Sammon asked 75 executives to check 100 recent records each from their own departments. On average 47% of newly created records had at least one critical error, and only 3% of the quality scores could be rated acceptable on the loosest standard.2

Google researchers then showed what that does downstream. Interviewing 53 practitioners building AI in high-stakes settings across India, the US and Africa, they found 92% had experienced at least one “data cascade”: a data problem that compounds invisibly through the system and surfaces much later as a failure. The paper’s title is a quote from one of them: “Everyone wants to do the model work, not the data work.”3

Gartner now predicts that through 2026, organisations will abandon 60% of AI projects that are not supported by AI-ready data. In its survey of 1,203 data management leaders, 63% either lacked the right data management practices for AI or were not sure.4

Layer two: pipelines

Google’s engineers made the canonical observation in 2015: only a small fraction of a real-world machine learning system is the ML code. The rest is data collection, verification, feature extraction, serving and monitoring, and that surrounding plumbing is where the long-run cost and fragility accumulate.5

RAND’s interviews with 65 experienced data scientists and engineers reached the same place from the failure side. Two of its five root causes are infrastructure: organisations lack the data needed to train an effective model, and they under-invest in the systems to manage data and deploy models.6

Layer three: the model in the real world

A model’s test score is measured with the other layers held perfect. Deployment removes that assumption.

Epic’s sepsis model was running in hundreds of US hospitals, with the vendor’s documentation citing an AUC of 0.76 to 0.83. When researchers at Michigan Medicine validated it on about 38,500 hospitalisations, the AUC was 0.63. It missed 67% of the patients who developed sepsis while firing alerts on 18% of all hospitalised patients.7 The demo number was real. The system around it was different.

Zillow had years of data and a well-known pricing model. Its home-buying arm bought 9,680 homes in one quarter of 2021, sold 3,032, lost about $304 million and was shut, with a quarter of the company’s staff cut. The CEO’s explanation: “the unpredictability in forecasting home prices far exceeds what we anticipated.”8 A model that was good enough to publish estimates was not good enough to commit capital on.

Layer four: adoption

The last multiplier is whether people act on the output, and it is the least forgiving. Dietvorst, Simmons and Massey showed across five experiments that people abandon an algorithm after watching it make a mistake, even when it is visibly outperforming the human alternative. Confidence is lost faster in an algorithm than in a person making the same error.9

This is why “good enough” compounds into “nobody trusts it”. Every upstream flaw eventually appears to a user as a wrong answer. A handful of visible errors is enough to send people back to the spreadsheet, at which point adoption is zero and so is the product. The Epic model’s alert rate is the clinical version: an alert on nearly one patient in five teaches staff to ignore alerts.7

Layer “Good enough” looks like What it does to the next layer
Data 47% of new records with a critical error2 Cascades surface late, far from the cause3
Pipelines Notebook to production by hand Model sees different data in production than in testing
Model AUC 0.76–0.83 on paper, 0.63 in use7 Wrong answers reach the user
Adoption Launched, not embedded Users see errors and stop using it9

Why the demo misleads

A demo tests one layer with the other three set to perfect: curated data, a hand-run pipeline, a motivated audience. It is an honest measure of the model and says nothing about the product.

MIT’s 2025 study counted the gap: 60% of organisations evaluated an enterprise-grade AI tool, 20% piloted one, 5% got one into production.10 RAND puts overall AI project failure above 80%, twice the rate for other IT projects.6

The spending pattern makes it worse. The model is the visible layer, so it gets the budget and the attention. In an O-ring process the return is set by the weakest layer, which is usually the one nobody wanted to fund.

What follows

  • Fund the weakest layer first. Extra model accuracy is worth little while data or adoption is the constraint.
  • Narrow the scope until every layer can be done properly. One use case with all four layers right beats ten with two.
  • Judge pilots end to end. Real data, real pipeline, real users, or it is still a demo.
  • Protect trust early. Users forgive an algorithm less than a colleague, so the first visible errors cost the most.

Half-hearted AI has a predictable result: four reasonable efforts, one unusable product, and a conclusion that the technology was oversold. The technology was fine. The arithmetic was wrong.

References

  1. M. Kremer, “The O-Ring Theory of Economic Development,” Quarterly Journal of Economics 108(3), 1993, pp. 551–575. Link
  2. T. Nagle, T. C. Redman and D. Sammon, “Only 3% of Companies’ Data Meets Basic Quality Standards,” Harvard Business Review, 11 September 2017. Link
  3. N. Sambasivan et al., “‘Everyone wants to do the model work, not the data work’: Data Cascades in High-Stakes AI,” CHI 2021. PDF
  4. Gartner, “Lack of AI-Ready Data Puts AI Projects at Risk,” 26 February 2025. Link
  5. D. Sculley et al., “Hidden Technical Debt in Machine Learning Systems,” NeurIPS 2015. Link
  6. J. Ryseff, B. F. De Bruhl and S. J. Newberry, The Root Causes of Failure for Artificial Intelligence Projects and How They Can Succeed, RAND, August 2024. Link
  7. A. Wong et al., “External Validation of a Widely Implemented Proprietary Sepsis Prediction Model in Hospitalized Patients,” JAMA Internal Medicine, 2021; figures as reported by Fierce Healthcare. Paper · Report
  8. NPR, “Zillow will stop buying and renovating homes and cut 25% of its workforce,” 3 November 2021. Link
  9. B. J. Dietvorst, J. P. Simmons and C. Massey, “Algorithm Aversion: People Erroneously Avoid Algorithms After Seeing Them Err,” Journal of Experimental Psychology: General 144(1), 2015. SSRN
  10. MIT NANDA, The GenAI Divide: State of AI in Business 2025. PDF
  11. E. Brynjolfsson, D. Li and L. Raymond, “Generative AI at Work,” NBER Working Paper 31161 (2023). NBER Digest