top of page

The Ferrari Engine in the Horse Cart: Why 95% of Enterprise AI Pilots Were Designed to Fail

Jul 13
5 min read
The Ferrari Engine in the Horse Cart

MIT's NANDA initiative put a number on what most executives already suspected: 95% of enterprise GenAI pilots deliver no measurable P&L impact. Billions invested; 5% with anything to show for it. The figure went viral and the methodology took some fire: the sample is modest, the framing is stark, and you're welcome to argue with the decimal point. But two things in that research are not in dispute, and they matter far more than the headline. First, the direction: high adoption, low transformation, across nearly every sector. Second, the diagnosis: the researchers blame integration and the organisation's inability to learn, not the models.

Everyone quotes the number. Almost nobody explains the structure behind it. The structure is this: the pilots didn't fail because the technology disappointed, and not because "change management" was underfunded. They failed because of their design. The failure was in the blueprint before the first line of the budget was signed.


The recipe everyone follows


The dominant pattern of enterprise AI adoption fits in one sentence. Take an existing process; find the step where a human does something repetitive; replace that step with an LLM.

It sounds sensible. It demos brilliantly. And it preserves, untouched, every assumption of the legacy process, from a data model designed for human eyes to a human role fixed by a division of labour made decades ago. All of it stays. One component changes. The result is a marginally cheaper version of exactly the same capability.

I call it the Ferrari engine in the horse cart. The engine is genuinely magnificent: that is what makes the failure so expensive and so confusing. But the cart's wheels weren't built for the torque, the axle transmits none of the power, and the driver is still holding reins. The engine's capability isn't the constraint and never was. So the pilot "succeeds" technically, delivers nothing commercially, and everyone concludes the engine was overhyped.

That's not transformation. That's an expensive way to stand still, and it's the single pattern behind most of MIT's 95%.


Why the failure is structural, not incidental


The recipe fails by design, not by execution, and the reason becomes visible the moment you ask what the swapped-in model is actually embedded in.

Execution without perception is blind. The pilot automates a task, but nobody has asked whether the task is the right one. The process being accelerated was designed for a market condition that may no longer exist; the pilot inherits that judgement without examining it. An agent performing the wrong work with precision is not an achievement.

Before an enterprise automates its response, it needs the capacity to perceive what the market is actually asking for: continuously, not once a year in the planning cycle. Almost no pilot portfolio contains a single project of that kind.

Execution without design is fragile. An LLM is a probabilistic system: the same input can produce different outputs, confidence varies, and the confident wrong answer is a design property, not a defect. The legacy process it's wired into is deterministic: it assumes the previous step is reliable, passes results downstream without questioning them, and produces data as unstructured exhaust.

Probabilistic logic inserted into a deterministic frame gives you the worst of both: the rigidity of the old process and the uncertainty of the new component, with nothing designed to absorb the mismatch. And because the surrounding process was never built to capture structured data, the pilot can't even learn from its own operation. It's frozen on day one, drifting out of alignment with the business at exactly the rate the business changes.

Notice that neither failure has anything to do with model quality. You could double the model's capability tomorrow and both problems would remain fully intact. That's why the fix isn't a better engine. There is no engine good enough to move a cart that wasn't built for one.


The other question


AI-Native-by-Design starts from a different question. Not "where can we insert AI into what we have?" but: if we were building this process today, knowing what AI can and cannot do, what would it look like?

Those two questions produce structurally different enterprises. The first treats AI as a component to be procured; the second treats it as a reason to redesign. The redesign, in my experience and in the book, rests on four principles. At the principle level:

Data as the primary output. In a legacy process, data is exhaust: inconsistent, unstructured, trapped in whatever format the original system produced. In an AI-Native process, every transaction is engineered to produce clean, structured, AI-ready data as its primary output. Not as a preference; as a prerequisite, because structured data is the connective tissue between AI components. Without it, your "system" is a collection of isolated pilots pretending to be one.

Probabilistic workflows. Legacy processes encode answers: if X, then Y. AI-Native processes branch on confidence: above the threshold, the decision proceeds; below it, it routes to a human. The firm stops encoding the answer into its processes and instead encodes the mechanism for finding the answer under whatever conditions prevail. That is what lets the process survive conditions its designers never anticipated.

The human in the loop as a design feature. Not a brake, and not a temporary concession to compliance. The human at the confidence boundary is decision-maker, trainer, and stabiliser at once, and every intervention generates labelled training data that makes the system more capable and more unique to this firm. Which leads directly to the fourth principle.

Continuous feedback loops. If the system isn't measurably better this month than last month from its own operation, it isn't AI-Native. It's a static tool with a subscription fee. The first three principles all exist to generate learning; this one closes the circuit.


None of these principles can be retrofitted into the horse cart, because each one contradicts an assumption the cart is built on. That's the honest answer to why 95% fail: they were attempts to get AI-Native results out of a legacy design, and the design won. It always does.

A case from the book gives the scale of the stakes: the same UK compliance requirement, met by two designs. The legacy build ran at roughly £150,000 a month; the AI-Native build at roughly £34,000. The gap didn't come from a smarter model. It came from a smarter system: data born structured, confidence thresholds in the process logic, human expertise placed at the boundary of uncertainty. Cost reduction without a demand thesis is arithmetic; inside an AI-Native architecture, it's strategy.


A diagnostic for your own portfolio

Before your next AI steering committee, put these questions to your current pilots. Did any of them begin with a blank sheet rather than an existing process? Does any produce structured data as its primary output, or do they all emit exhaust? Does any route decisions on confidence, or do they all assume the model is right? Is any of them measurably better this quarter than last, from its own operation? And is there a single project in the portfolio whose job is perceiving the market rather than executing a task?

If the answers are mostly no, your portfolio is a menu of ways to spend money: a ranked list of Execution projects with no perception and no design behind them. That's not a criticism of your team. It's the industry default. The point of asking is that the default is now measurably, publicly, 95% likely to fail, and the alternative is an architecture, not a purchase.

That architecture, its three layers and what it demands of the firm, is the subject of the final chapters of Architecture of Intellect, which also gives the people accountable for all of this an honest, vendor-free account of the engine itself. It's available now at notesofea.com/book, and the next article on this site will set out the three-layer architecture in full.

Igor Ageyev is an enterprise architect with more than thirty years in banking and financial services technology. His AI experience spans two eras: expert systems in the 1990s, and LLM integration in enterprise operating models since 2021. He writes at notesofea.com.

Related Posts

See All
The revolution is becoming a framework

Each generation of AI models moves from a stochastic toy toward a controllable framework. That's not disappointment; it's when the technology becomes fit for real work inside a process.

 
 
 

Comments

Rated 0 out of 5 stars.
No ratings yet

Add a rating

Get the next argument first. Articles, campaign posts and book news. No more than one email a week.

bottom of page