
Every business conversation about AI eventually turns to data. Usually in the wrong direction.
Leadership teams ask how much data they have, whether their data lake is big enough, or whether they need to buy more of it. Those are the wrong questions. The organisations getting real value from AI aren't the ones with the most data. They're the ones with data they can actually trust, structure, and use.
Great AI needs great data. Not more of it. That distinction is the difference between an AI initiative that pays for itself and one that quietly stalls in a pilot nobody scales.
It's also the reason so many AI projects start with the wrong first move. Businesses buy a platform, commission a model, or greenlight a pilot before anyone has properly asked whether the information behind it is accurate, structured, or even findable. The AI strategy gets built first. The data strategy gets bolted on afterwards, if at all. That ordering is backwards, and it's expensive to reverse once a project is already underway.
It's tempting to assume that AI progress is a technology story: better models, faster chips, smarter agents. The evidence points elsewhere. McKinsey's research on enterprise AI adoption found that while AI use has become close to universal across businesses, only around a third of organisations have managed to scale it beyond pilots and isolated use cases. The rest stay stuck running experiments that never become operational systems.
The pattern behind that gap is consistent. Businesses that scale successfully treat clean, integrated, well-governed data as a prerequisite, not an afterthought. Businesses that stall usually have the opposite problem: fragmented systems, inconsistent records, and no one accountable for fixing either.
The model is rarely the constraint. The data underneath it is.
It's easy to see why volume feels reassuring. The world is generating an extraordinary amount of information. Statista's data on global data creation puts the total volume of data created, captured and consumed worldwide at around 149 zettabytes in 2024, with that figure forecast to climb further still. There is no shortage of raw material.
What there is a shortage of is usable material. Most of that growth is in unstructured formats: documents, emails, transcripts, PDFs, images, sensor logs. According to Gartner, the vast majority of enterprise data falls into this category, and it grows faster than structured data every year. Much of it becomes what IBM calls "dark data": information a business collects and stores but never actually puts to use.
Feeding an AI system more of this kind of data doesn't make it smarter. It makes the underlying problem bigger. A model trained on inconsistent, duplicated or poorly labelled information doesn't hedge its answers. It produces confident output that looks correct and often isn't, which is a harder problem to catch than an obviously wrong one.
Volume was never the differentiator. Quality, structure and control over that data are.

Businesses that get value from AI tend to share a few concrete traits in how they handle their data, well before any model gets built.
It's accurate and current. Records reflect reality, not a snapshot from eighteen months ago. Duplicate entries, dead fields and unreconciled figures are found and fixed as a matter of routine, not discovered when a report looks wrong.
It's structured for use, not just storage. Information sitting in disconnected spreadsheets, legacy systems and departmental silos might technically exist, but it isn't usable at scale. Clear schemas and consistent definitions are what let a system interpret data reliably instead of guessing at it.
It's accessible without a manual workaround. If answering a straightforward business question means exporting data into a spreadsheet and stitching it together by hand, that data isn't ready to support AI, no matter how much of it there is.
It's governed continuously, not cleaned once. Data quality decays the moment nobody owns it. A one-off clean-up buys a business a few tidy months at best, unless clear ownership, monitoring and access controls keep it that way.
None of this is exotic. It's also not optional. It's the groundwork that determines whether an AI investment compounds or quietly evaporates.
Poor data doesn't just slow an AI project down. It changes what that project can honestly deliver.
A forecasting model trained on sales figures that three regional teams record differently won't produce a forecast anyone trusts, it will produce three teams' worth of disagreement dressed up as a single confident number. A customer service assistant built on outdated product records will answer questions accurately right up until the moment the product range changes and nobody updated the source. A reporting tool that pulls from five systems with five different definitions of "active customer" won't resolve that inconsistency. It will just make it faster to distribute.
None of these are AI failures in the sense of the technology not working. The model does exactly what it was built to do. The problem sits one layer down, in the data strategy that was never put in place before the build started. That's a costly lesson to learn after the investment has already been made, and a straightforward one to avoid by addressing data quality and structure first.

Deloitte's most recent State of AI in the Enterprise research breaks organisational AI readiness down across four dimensions: governance, technical infrastructure, data management and talent. Data management and governance both score among the weakest of the four, with businesses reporting considerably more maturity in infrastructure than in the data discipline needed to run on top of it.
That's a telling gap. Businesses are willing to invest in compute, platforms and tooling. Far fewer have done the less visible work of getting their own data estate into a state that can actually support what they're building. The technology is often ready before the organisation is.
None of this requires solving every data problem a business has ever accumulated. It requires an honest, specific starting point.
The organisations pulling ahead with AI aren't necessarily working with more data than their competitors. They're working with data they can trust, structure and put in front of a model with confidence. That's a less dramatic story than the latest AI headline, but it's the one that actually determines whether an investment pays off.
Get the data foundation right, and the benefits show up well before the first model reaches production: cleaner reporting, faster decisions, less manual reconciliation. Get it wrong, and no amount of algorithmic sophistication will make up the difference. The businesses that treat data quality, structure, accessibility and governance as the starting point, rather than a box to tick after the AI has already been bought, are the ones whose AI investment actually compounds instead of stalling out.
If your business is weighing up an AI investment and isn't confident the data behind it can support it, our data consulting team can help assess where the real gaps sit and build the trusted, governed data foundations that make AI investment worth making in the first place.
Have a project in mind? No need to be shy, drop us a note and tell us how we can help realise your vision.
