AI Data Readiness
Deploying AI onto a weak data foundation does not produce a smaller benefit. It can produce a negative one. Here is what readiness actually means and how to scope it so the work finishes.
When I took the interim CTO seat at a luxury destination-club and vacation-rental operator, the ask on the table was AI. What the business actually needed first was for its pricing data to be correct.
That sequencing is not a detail. It is the whole story of why most AI programs stall, and it is the single most useful thing I can tell an operator who is about to spend real money on this.
You have seen it. 95% of enterprise AI projects fail. It comes from a 2025 MIT Media Lab report, and it has been repeated so widely that it now functions as common knowledge. It should not.
The research reviewed 300+ publicly disclosed initiatives, interviewed 52 organizations, and surveyed 153 leaders — recruited at four industry conferences, over a six-month window. The 5% figure refers to a subset: task-specific generative AI tools that reached production inside that window. The report's own limitations section acknowledges selection bias and says the window "may be insufficient to fully assess successful deployment." Wharton's Kevin Werbach, after several readings, said he could find no further support in the document for the 95% claim.
I raise this not to be pedantic but because it matters how boards are being informed. If you are going to make a capital allocation decision on the strength of a number, the number should survive being looked at. There are better ones.
S&P Global surveyed over a thousand organizations across North America and Europe and found the share scrapping most of their AI initiatives rose from 17% in 2024 to 42% in 2025, with the average organization abandoning 46% of proofs of concept before production. That is a clean year-over-year comparison with an unambiguous definition. It tells you something real, and it is bad enough on its own.
Gartner surveyed more than 1,200 data management leaders and expects 60% of AI projects to be abandoned through 2026 for want of AI-ready data — with 63% of organizations lacking, or unsure they have, the data management practices required.
The most rigorous framing I've found comes from Google's 2025 DORA research, which surveyed roughly 5,000 technology professionals. Their central finding is worth committing to memory:
"AI doesn't fix a team; it amplifies what's already there."
DORA found AI adoption correlates positively with delivery throughput and negatively with delivery stability. It accelerates the front of the pipeline and exposes everything weak downstream. They name "healthy data ecosystems" as the capability that determines whether AI's effect on organizational performance is positive at all.
That is not a caveat. It is a sign change. Deploying AI onto a poor data foundation does not produce a smaller benefit — it can produce a negative one, because you have industrialized the propagation of wrong answers.
"Data readiness" gets used as though everyone agrees what it covers. In practice I assess six things, and I'd encourage any board to ask about them in exactly these terms:
At the destination-club operator, the pricing data had degraded to the point where it was actively suppressing conversion. Availability and rates disagreed across distribution channels. There was no reporting layer anyone trusted. Integration debt had accumulated across three separate channel and property-management platforms.
Had we started with agents, we would have built something that confidently made decisions on wrong inputs and did it faster than a human could catch. So we built the data lakehouse first — dashboards, automated data-quality monitoring, one trusted view of pricing and inventory. Only then did the pricing, data-quality, and conversion agents go onto AWS Bedrock.
Those agents are in production against live revenue data today. They work because the layer beneath them is sound. The order of operations was the entire intervention.
There's a reasonable objection: data readiness can become an excuse for permanent delay. Every data program can be extended indefinitely, and "we're not ready yet" is a comfortable place for a cautious organization to live.
I take that seriously. The answer is not to sequence the whole enterprise's data estate before doing anything — it's to sequence the specific data the first use case depends on. At the club operator, we did not fix all the data. We fixed pricing and availability, because that's what the first agents would touch. Scope readiness to the use case, ship, then widen.
The uncomfortable version of all this: if your AI pilot failed, the pilot probably wasn't the problem. Something underneath it was, and it was there before the pilot started. That's actually good news — it means the fix is knowable, and you were going to need it regardless.
The Production AI Blueprint is a fixed-fee engagement that sorts out your data, governance, controls, and ownership — and hands you a sequenced roadmap you own outright.