Most AI pilots never get cancelled. They demo well in the spring, pick up a second round of scope over the summer, and by autumn the two people who built the thing have been pulled onto something with a firmer deadline, without anyone writing the word “failed” anywhere.
The question of why AI pilots fail to scale has an answer upstream of the model, in the gap between the curated extract the pilot was built on and the data estate it had to survive in.
BARC’s 2026 research puts AI readiness rather than model quality at the center of that gap, and this article covers what the evidence shows, why capable teams skip the step anyway, and what to do about it this quarter. We wrote it for mid-market data teams, big enough that an AI initiative has real money attached, not so big that there is a dedicated discovery function to hand the work to.
At a Glance
- In BARC’s 2026 survey of 225 data and AI leaders, 70% of organizations reported that less than half of their unstructured data is discoverable or usable for AI.
- BARC still classes only 23% of organizations as AI Leaders, against 20 to 21% in the two prior surveys.
- What we see with mid-market teams is that changing the model almost never restarts a stalled pilot, because the breakage happens further upstream.
- Classification and validation lead BARC’s list of preparation activities at AI-mature organizations, at 60% and 56%. Vectorization and fine-tuning do not come close.
- Discovery gets skipped for rational reasons: no demo at the end of it, and a coordination bill spread across departments that do not report to you.
- A smaller estate is a faster estate to map, which is where the mid-market advantage sits.
- Start with a two-column inventory: the data you hold, against the data you can use.
Who This Is For:

The Pilot Did Not Stall on the Model
Think about how these projects get built. Someone pulls a clean extract, a few thousand CRM rows plus a folder of documents already tidied for another project, and the model performs well against it. Then the system meets the real environment, where the same customer exists three times under two spellings, half the contracts sit in a SharePoint nobody has audited since the migration, and a file server labeled temporary in 2020 is still load-bearing. Performance falls apart at precisely the moment the thing stops being a demo.
What Stalling Actually Looks Like
It is rarely a decision. The next milestone gets pushed for a reason that sounds procedural, the champion’s attention moves to a budget cycle, and the environment quietly stops being maintained. Six months later somebody asks whether anything came of that AI project. Nothing broke loudly enough to earn a post-mortem, which is what makes it hard to diagnose.
Why the Model Gets Blamed First
Ask a room why AI projects fail and the model comes up first, because it is the visible component and the cheapest thing to change. So teams swap models, run the same evaluation against the same fragmented inputs, and get a variation on the same result. Our reading is that the model draws the blame because it is the only part of the stack you can point at without implicating another department.
What BARC’s 2026 Research Found

The number that reframes the conversation comes from Unstructured Data for AI Innovation, the study Merv Adrian and Kevin Petrie presented at BARC’s 2026 Denver data retreat. Across 225 data and AI leaders, 70% of organizations reported that less than half of their unstructured data is discoverable or usable for AI. The documents, emails, contracts, and meeting notes carrying most of your institutional knowledge are effectively invisible to the systems being built on top of them.
Set that against the wider picture. Of the organizations BARC surveyed in 2026, only 23% qualify as AI Leaders, defined as executive-led, governed, and production-deployed. That share has barely moved in three years, sitting at 20 to 21% in the two prior surveys. Our reading of a line that flat is that the constraint has never been the models.
Unstructured Data Is the Context Layer
Structured data tells you what happened. Unstructured data tells you why, which is the difference between a system reporting last quarter’s churn number and one that can tell you the three accounts you lost all raised the same onboarding problem in their tickets. Strip out that context layer and the output stays precise while getting shallower, and a confident shallow answer is much harder to catch than an obviously wrong one. Whether the answers come back grounded or merely fluent is settled by a solid data foundation rather than by the model.
Classification and Validation Are What Mature Teams Actually Do
This is where the research gets specific about how to prepare data for AI. BARC found classification at 60% and validation at 56% as the most widely adopted and consistently effective preparation activities among AI-mature organizations. Not vectorization. Not fine-tuning. Knowing what a piece of data is, and whether it can be trusted. Neither needs a new vendor, and together they cover most of the distance between a pilot and something a business unit can rely on.
Why Data Discovery Keeps Getting Skipped
Nobody skips this out of carelessness. Discovery has no demo at the end of it, so no stakeholder ever sees something and gets excited. It needs coordination between IT, legal, operations, and whoever owns the content, and those people have their own quarters to survive. Under pressure to show results by somebody else’s date, building against the structured data already available is reasonable enough. The error is one of category rather than effort: discovery gets treated as pre-work when it is the foundation.
The Governance Exposure in the Same Gap
The same undiscovered data creates a second problem that surfaces later, at higher cost. BARC’s research shows 33% of organizations have inconsistent or no controls for data bias, and 28% lack lineage tracking. The damage from those two gaps compounds fastest as organizations move toward AI agents that act with less human supervision. You cannot govern what you have not mapped, so discovery is the entry point for any governance posture.
What the Organizations That Get Through Do Differently
BARC frames the payoff of doing this early as the rework it prevents once a system is in production. Our reading of what that rework looks like from the inside: output that is confidently wrong, trust that drains within weeks, and a tool nobody opens. Trust is far cheaper to establish than to rebuild.
Two characteristics show up repeatedly in the organizations BARC identifies as succeeding, and neither is a technology choice.
They Built an Inventory First
They started with a map. Where does our data live, what form is it in, who owns it, how sensitive is it, how current is it, and can we find it again reliably. Classification and validation happened on top of that inventory, before anything was connected to an AI workflow, which is the practice those 60% and 56% figures point to. Notice how ordinary it is. None of it requires a platform decision.
They Treat Human Review as Part of the System
BARC found that 41% of AI-mature organizations rely on human validation as a primary AI readiness measure. That is a design choice rather than an admission of failure. The goal is not keeping a person in the loop forever, it is making the checking deliberate instead of reactive, so review happens at defined points rather than whenever a number looks off.
Why Mid-Market Teams Are Better Placed Than They Think
An enterprise running this exercise coordinates across acquired subsidiaries with their own systems and a governance function that signs off in writing.
You have less data, fewer systems, and a much shorter distance between the person who knows where the contracts are and the person who needs them. An AI readiness assessment covers the same ground either way; what shrinks is the calendar rather than the work.
The advantage is speed rather than exemption, and skipping it because the estate is small produces the same stalled pilot on a smaller budget.
Five Things to Do This Quarter
| Action | How to Approach It |
|---|---|
| Run a discovery audit before the next AI initiative | Ask where the data lives, whether it is findable, whether it is current, and whether anyone has ever validated it |
| Separate data you have from data you can use | Build a two-column inventory: assets you possess against assets that are classified, validated, and accessible |
| Make classification a habit, not a project | Assign ongoing ownership at the team or system level, rotating rather than permanent |
| Start with the most-referenced content | Identify the 20% of unstructured content that shows up most often in decisions and make that discoverable first |
| Design for distributed data | BARC’s research shows on-premises and hybrid environments persist rather than consolidating, so build for data that stays where it is |
Two rows there carry more weight than the rest.
The two-column inventory is the one that changes internal conversations. Most teams already suspect the gap between what they hold and what they can use is wide, and what they lack is a way to show it. Two lists side by side turn a vague argument about whether the team is moving fast enough into a specific, fundable piece of scope. Nobody argues with a list.
Designing for distributed data contradicts advice most teams have heard for a decade. The single source of truth is a good principle and a poor project plan, and BARC’s finding that hybrid environments persist rather than consolidate suggests the consolidation-first sequence fails for structural reasons. Build your access layer assuming the data stays spread out, because the pipeline work that makes distributed data usable pays off in weeks rather than at the end of an unfunded migration.
How ClicData Fits
None of the work above requires buying anything, which is worth stating before we describe our own platform. Where a platform helps is holding the inventory together once it exists.
500+ connectors put the sources behind a decision in one place, so what your pilot ran against and what production reads are the same data, not two copies that drifted.
Transformation and validation inside Data Flows run before anything reaches a dashboard, making the checking systematic rather than something an analyst does by eye on a Friday.
Sources stay where they are while the analytics layer consolidates on top, fitting BARC’s finding on hybrid environments instead of fighting it.
Role-based access at the data layer holds regardless of which AI-powered analytics feature or user reaches it.
File Storage keeps unstructured files in that same governed environment, though holding a contract is not the same as knowing what it says.
The honest limit sits above the tooling. Software can catalog and classify, and it keeps getting better at both. What no tool settles is which of those six years of meeting notes bear on the decision you are about to automate, or who is allowed to act on them. That judgment comes from people outside the data team, and no purchase removes it.
The full research is in What Mid-Market Teams Get Wrong About AI.

Conclusion: Sequencing Is the Advantage
The 23% BARC identifies as AI Leaders are not running different models or outspending everyone else by a margin that explains the gap. They resolved discovery, quality, and governance before scaling AI, and their lead has held because sequencing is a discipline rather than a purchase.
So before the next AI initiative gets scoped, take the dataset it depends on most and build the two columns. What you hold, against what is classified, validated, and accessible. It either confirms you are ready or surfaces the problem while it is cheap to fix, and both beat another pilot that only demos well.
FAQs
Why do AI pilots fail to reach production?
Most of the time the model works and the data underneath it does not. A pilot runs against a curated extract; production runs against the real estate, which is fragmented, partially documented, and spread across systems with different owners. BARC’s 2026 research points to discoverability and data quality as the binding constraints, which is why swapping models rarely rescues a stalled project.
How much of a company’s data is usable for AI?
Less than most leadership teams assume. BARC’s 2026 research puts 70% of organizations below the halfway mark on how much of their unstructured data can be discovered or used, and documents, email, and contracts are where that shortfall concentrates. Businesses running on paperwork rather than transactions generally have the most ground to cover.
What is data discovery, and why does it come before AI?
Data discovery is the exercise of establishing where your data lives, what form it takes, who owns it, how sensitive it is, and whether it can be found again reliably. It comes first because everything downstream depends on it. You cannot classify what you have not found, cannot validate what you have not classified, and cannot govern any of it without the map.
Should we consolidate all our data before starting an AI project?
No, and BARC’s research suggests the consolidation-first sequence is part of why teams stall. On-premises and hybrid environments are persisting rather than converging, so a project that waits for a single source of truth usually waits indefinitely. Map what you have, connect it where it sits, and consolidate the analytics layer rather than the storage.
What is the difference between data classification and data validation?
Classification answers what a piece of data is: its type, sensitivity, owner, and where it belongs in your taxonomy. Validation answers whether it can be trusted: is it accurate, current, complete, and consistent with the other sources describing the same thing. BARC found classification at 60% and validation at 56% among AI-mature organizations, and they work as a pair rather than as alternatives.
How long does a data discovery audit take for a mid-market team?
Scoped to the sources behind one initiative rather than the whole estate, what we see with mid-market clients is a few weeks of work rather than a quarter. Scoped to everything, it becomes the project that never finishes. Start with the 20% of content your decisions reference most and extend once the method is proven.
Does better data quality fix AI hallucinations?
It removes one major cause and leaves the others in place. Clean, current, well-governed inputs eliminate the failures where a model confidently reports something because the underlying record was wrong or stale. What data quality cannot address is the model generating plausible text with no grounding at all, which is a property of how these systems work. This is why 41% of AI-mature organizations in BARC’s research still rely on human validation as a primary readiness measure, and why mature programs treat review as a permanent component rather than temporary scaffolding.


