AI Data Readiness

Your AI problem is a data problem wearing a costume

Deploying AI onto a weak data foundation does not produce a smaller benefit. It can produce a negative one. Here is what readiness actually means and how to scope it so the work finishes.

When I took the interim CTO seat at a luxury destination-club and vacation-rental operator, the ask on the table was AI. What the business actually needed first was for its pricing data to be correct.

That sequencing is not a detail. It is the whole story of why most AI programs stall, and it is the single most useful thing I can tell an operator who is about to spend real money on this.

First, about that 95% statistic

You have seen it. 95% of enterprise AI projects fail. It comes from a 2025 MIT Media Lab report, and it has been repeated so widely that it now functions as common knowledge. It should not.

What the study actually measured

The research reviewed 300+ publicly disclosed initiatives, interviewed 52 organizations, and surveyed 153 leaders — recruited at four industry conferences, over a six-month window. The 5% figure refers to a subset: task-specific generative AI tools that reached production inside that window. The report's own limitations section acknowledges selection bias and says the window "may be insufficient to fully assess successful deployment." Wharton's Kevin Werbach, after several readings, said he could find no further support in the document for the 95% claim.

I raise this not to be pedantic but because it matters how boards are being informed. If you are going to make a capital allocation decision on the strength of a number, the number should survive being looked at. There are better ones.

S&P Global surveyed over a thousand organizations across North America and Europe and found the share scrapping most of their AI initiatives rose from 17% in 2024 to 42% in 2025, with the average organization abandoning 46% of proofs of concept before production. That is a clean year-over-year comparison with an unambiguous definition. It tells you something real, and it is bad enough on its own.

The constraint is almost always the data

Gartner surveyed more than 1,200 data management leaders and expects 60% of AI projects to be abandoned through 2026 for want of AI-ready data — with 63% of organizations lacking, or unsure they have, the data management practices required.

The most rigorous framing I've found comes from Google's 2025 DORA research, which surveyed roughly 5,000 technology professionals. Their central finding is worth committing to memory:

"AI doesn't fix a team; it amplifies what's already there."

DORA found AI adoption correlates positively with delivery throughput and negatively with delivery stability. It accelerates the front of the pipeline and exposes everything weak downstream. They name "healthy data ecosystems" as the capability that determines whether AI's effect on organizational performance is positive at all.

That is not a caveat. It is a sign change. Deploying AI onto a poor data foundation does not produce a smaller benefit — it can produce a negative one, because you have industrialized the propagation of wrong answers.

What data readiness concretely means

"Data readiness" gets used as though everyone agrees what it covers. In practice I assess six things, and I'd encourage any board to ask about them in exactly these terms:

  • Representativeness. Gartner's definition is unusually good: the data must contain the patterns, errors, outliers and unexpected cases the model will actually meet. Clean training data that omits your messy real-world cases produces a model that fails precisely when it matters.
  • Ownership. Is one named executive accountable for data, analytics, and AI together? Recent MIT Sloan work puts the share of organizations with a chief AI officer or equivalent at 38%. Split accountability is how data quality becomes nobody's problem.
  • Definitions. Does "customer," "revenue," or "active" mean one thing across your systems? This is the least glamorous item on the list and the one that most often blocks everything else.
  • Lineage. Can you reconstruct which data produced a given output? Without it you cannot debug a bad answer, and you cannot answer a regulator or an auditor.
  • Freshness. Gartner's phrasing again: AI-ready data "is not one and done." Readiness decays. It is an operating discipline, not a project with an end date.
  • Access. Can teams and tools discover the data, with governance intact? DORA separates accessibility from quality for good reason — excellent data nobody can reach is functionally absent.

What this looked like in practice

At the destination-club operator, the pricing data had degraded to the point where it was actively suppressing conversion. Availability and rates disagreed across distribution channels. There was no reporting layer anyone trusted. Integration debt had accumulated across three separate channel and property-management platforms.

Had we started with agents, we would have built something that confidently made decisions on wrong inputs and did it faster than a human could catch. So we built the data lakehouse first — dashboards, automated data-quality monitoring, one trusted view of pricing and inventory. Only then did the pricing, data-quality, and conversion agents go onto AWS Bedrock.

Those agents are in production against live revenue data today. They work because the layer beneath them is sound. The order of operations was the entire intervention.

The honest counterargument

There's a reasonable objection: data readiness can become an excuse for permanent delay. Every data program can be extended indefinitely, and "we're not ready yet" is a comfortable place for a cautious organization to live.

I take that seriously. The answer is not to sequence the whole enterprise's data estate before doing anything — it's to sequence the specific data the first use case depends on. At the club operator, we did not fix all the data. We fixed pricing and availability, because that's what the first agents would touch. Scope readiness to the use case, ship, then widen.

What to do about it

  1. Pick the use case first, then assess only the data it touches. A general data readiness program has no end condition. A use-case-scoped one finishes.
  2. Run the six-question audit above on that slice. Representativeness, ownership, definitions, lineage, freshness, access. Anything that fails is scope for the phase before the AI phase.
  3. Name one accountable executive for data, analytics, and AI. Not a committee. If three people own it, no one does.
  4. Build the monitoring before the model. Automated data-quality checks on the inputs are what let you deploy an agent and sleep. Retrofitting them after an incident is considerably more expensive.
  5. Ask where your unofficial AI usage already is. The MIT report's genuinely useful finding was that employees at 90%+ of surveyed companies use personal AI tools for work while only 40% of the companies had bought a subscription. Your data is already going somewhere. Better to know where.

The uncomfortable version of all this: if your AI pilot failed, the pilot probably wasn't the problem. Something underneath it was, and it was there before the pilot started. That's actually good news — it means the fix is knowable, and you were going to need it regardless.

Sources

  1. MIT Media Lab Project NANDA, The GenAI Divide: State of AI in Business 2025
  2. Futuriom, Why we don’t believe MIT NANDA’s AI study (August 2025), including comment from Wharton’s Kevin Werbach
  3. S&P Global via CIO Dive, AI project abandonment rises (n=1,000+)
  4. Gartner, Lack of AI-ready data puts AI projects at risk
  5. Google Cloud, 2025 DORA report and Healthy data ecosystems
  6. MIT Sloan, Action items for AI decision makers, 2026

Kurt Wysock

Fractional & interim CTO and board technology advisor. Currently interim CTO at a luxury destination-club operator, head of the Technology & Architecture Office at a global consulting firm, and founder of JSummit Consulting. 30+ years in enterprise architecture across energy, hospitality, manufacturing, financial services, retail, and healthcare. TOGAF® 9 Certified. More about Kurt · Get in touch

← Cloud repatriationWhere agents actually work →

Want this applied to your systems?

The Production AI Blueprint is a fixed-fee engagement that sorts out your data, governance, controls, and ownership — and hands you a sequenced roadmap you own outright.