BLACKSIG SYSTEMS/Resources/Is Our Data Good Enough for AI Automation

// Resources

Is Our Data Good Enough for AI Automation?

The objection that stops more automation projects at the first meeting than cost does. Which data problems actually stop a build, which ones are noise, how to check your own in an afternoon, and what to fix first.

Last updated · 2026-09-29

The short version

Probably, for the first thing you automate. The question that matters is not whether your data is good in general, it is whether the records one specific process touches are correct and reachable. A business with a messy shared drive, four spreadsheets and a CRM nobody trusts can still automate quoting, invoice coding or after hours booking, as long as the handful of records those processes use are right. What actually stops a build is narrower than most owners fear: one entity stored under several identities, a system nothing can write to, rules nobody can state, and inputs so inconsistent there is nothing stable to build against. The industry gets this wrong in both directions. Gartner named poor data quality first among the reasons it expected at least 30% of generative AI projects to be abandoned after proof of concept by the end of 2025, while plenty of firms will sell you a data platform you do not need before your first automation runs.

What does AI ready data actually mean?

AI ready data means the records a specific process depends on are accurate, complete enough for that process, and reachable by software. It does not mean a warehouse, a lake or a governance framework.

Four properties matter, and they are all about a particular process rather than about the company. Accuracy, meaning the record matches reality: the customer's current address, the supplier's real bank details, the actual price on that contract. Consistency, meaning one real thing is stored as one record, not three. Reachability, meaning software can read and write it through an API, an export or a supported integration. And enough history for the task at hand, which for most process automation is very little, because the system is applying rules rather than learning patterns.

History is where the confusion starts. Prediction needs history: forecasting demand or churn wants years of clean records. Process automation mostly does not. A system that reads an invoice, matches it to an order and codes it needs today's supplier list and today's chart of accounts, not five years of tidy accounting. Owners who have been told they need a data program before they can automate anything have usually been quoted for the prediction case when they asked about the process one.

The honest test is per process, and it takes an afternoon. Name the process. List every record it reads or writes. For each one, ask: is it correct today, does one real world thing map to one record, and can software reach it. Anything that fails is a specific piece of work with a specific size, which is far more useful than a general verdict about whether the company is data mature.

The World Economic Forum's January 2026 piece on data readiness reports that fewer than one in five organisations consider themselves data ready, and that more than half of business leaders cite data quality and availability as a major challenge to accelerating AI adoption. If company wide readiness were the requirement, almost nothing would ever ship. Scoped per process, plenty does.

Which data problems actually stop an AI automation build?

Four problems stop a build. Everything else on the usual readiness checklist is either a cleanup task that runs alongside the work or an enterprise concern that does not apply to you yet.

Problem Why it stops the build What fixing it involves
One real thing stored as several recordsThe system cannot tell which customer, supplier or job it is looking at, so every output is suspectDeduplication and a decision about which record wins. Usually days, not months
A system nothing can write toThe automation produces an answer and a person re-keys it, which removes the pointCheck for an API or supported integration first. Where there is none, the process has to be redesigned around what can be reached
Rules that exist only in somebody's headNothing can be built until the rules are stated, because the exceptions are the actual specificationSit with the person who does the work and write down what happens when the input is wrong
Inputs with no stable shapeIf every instance of the input differs and no pattern holds, there is nothing to build againstSample fifty real examples. If they cluster, build. If they genuinely do not, pick a different process first
Missing historyOnly blocks forecasting and prediction, not process automationStart with process work and revisit prediction later
No data warehouseDoes not block process automation at allIgnore until something actually needs it
No formal governance frameworkDoes not block a first build. Access control and an audit trail do matter, and both are narrower than a frameworkDecide who can authorise system access to which system, in writing

The first row is the one we hit most often, and the one most often mistaken for a technology problem. Three spellings of the same customer, two supplier records with different bank details, a job number that means something different in the field app than in the ledger. None of that is exotic and all of it is fixable, but it has to be fixed before a system relies on it rather than after somebody notices a wrong output.

The third row is the most underestimated. A process that runs fine on human judgement has rules, they are just undocumented, and they live with whoever has been doing the job for nine years. That person is the specification. Getting the rules out of their head is the single highest value hour in a scoping engagement, and it is why we insist the people who do the work are in the room rather than only the people who manage them.

Do you need to clean your data before you automate anything?

Clean the records the first automation touches. Leave the rest alone until something needs them.

The instinct to fix everything first is understandable and it is how a six week project becomes an eighteen month one. Company wide data cleanup has no natural end, no visible benefit while it is happening, and a habit of stalling when the person driving it gets busy. Meanwhile the process that was costing you twenty hours a week is still costing you twenty hours a week.

Scoping the cleanup to one process changes the size of the job entirely. Automating invoice coding needs a correct supplier list and a chart of accounts that reflects how you think about costs. It does not need your customer records fixed. Automating after hours booking needs the service list, the calendar and the customer record. It does not need the supplier list. Each build brings its own small cleanup, and after three or four builds a surprising amount of the company's data is correct as a side effect, with each piece of tidying paid for by something that started working.

There is real evidence that the general problem is widespread. Validity's State of CRM Data Management in 2026 report, from a survey of 500 marketing professionals across five countries, found 62% of organisations losing revenue directly because of poor CRM data quality, and nearly a third of teams spending six or more hours a week fixing and reconciling data. That is an argument for fixing data. It is not an argument for fixing all of it first, because the organisations in that survey had the same problem before AI and will have it afterwards.

One exception is worth naming. Where the bad data is what you would be automating against for the decision itself, cleanup comes first and is not negotiable. A pricing system driven by a price list nobody has maintained will confidently apply wrong prices faster than a person would. When we find that, we say the first engagement is a data one, even when it is not the work we were asked for.

How do you check whether your data is ready, in an afternoon?

Run four checks on one process. Each is a spreadsheet exercise rather than a technical audit, and together they give a clearer answer than a readiness assessment.

Count the duplicates. Export the records the process uses, sort by name, look for the same entity twice. Do the same by email, phone or account number. The number you find tells you the size of the deduplication job, and finding almost none tells you something useful too.

Trace one item end to end. Pick a real recent case, one invoice, one enquiry, one job, and follow it through every system it touched. Note where somebody typed the same information twice. Every re-keying point is either the automation's job or the reason the automation cannot reach that far, and you will find at least one you did not know about.

Ask the person doing the work what happens when the input is wrong. Not the manager. The person. Their answer is the exception list, and the exception list is the specification. If they cannot answer because it depends, that dependency is what the build has to handle and it needs writing down.

Check what can be written to. For every system in the chain, find out whether it has an API, a supported integration, or an export and import a machine can drive. This is the check that most often changes the plan, because a system that can only be read forces the process to be redesigned rather than automated in place. The practical version is in does AI automation work with my software.

Those four produce something more useful than a score: a named list of fixes with sizes attached. That is also roughly what the first stage of our own work produces, alongside the decision about which process is worth doing first. What that engagement covers is in what an AI audit is, and what comes out of it is in what an AI roadmap should include.

What if most of our information is in documents, emails and people's heads?

Documents and email are workable inputs, and the heads are the part that needs attention.

Reading documents is settled work. Invoices, contracts, forms, reports, scanned pages, email threads: extracting fields and meaning from these is proven rather than experimental, which is why so much of what we build starts with a document arriving. A business whose records are mostly unstructured is not behind. It is in the ordinary case, and often a better candidate than a business with a half configured data platform, because there is less to unpick.

Email is workable with one caveat. A thread contains the decision, the context and four irrelevant replies, so a system reading it needs a clear definition of what it is looking for. That definition is a scoping question rather than a technical limit.

Knowledge in people's heads is the real constraint, and not because it cannot be captured. It is that nobody has written it down and the person holding it is busy. A quoting process where the experienced estimator adjusts for site access, a triage process where the senior nurse knows which symptoms need a call back, a coding process where the bookkeeper knows which supplier invoices always arrive misdescribed: the judgement is the value and it is undocumented. Extracting it takes that person's time, and the project has to budget for it rather than hope.

What none of this requires is that you have been diligent for the last five years. MIT's Project NANDA report, The GenAI Divide: State of AI in Business 2025, attributes pilot failure to brittle workflows, systems that retain no context, and poor fit with how work actually runs, rather than to insufficient data. Fit is a design problem. Design is what the strategy stage is for, and where the maturity of each capability is honest rather than aspirational, which is why the industry boards mark every item proven, working or early.

Does better data pay for itself before the automation does?

Sometimes better data does pay for itself first, and it is worth checking, because the cleanup often has a return of its own that nobody counted.

Three returns show up before any automation runs. Duplicate records cost money directly: duplicate invoices get paid, duplicate outreach annoys customers, and duplicate job records mean two people drive to the same address. Bad contact data costs conversions, which is the finding underneath the Validity result above, where 62% of organisations reported losing revenue directly because of poor CRM data quality. And a correct price list or rate card stops the quiet margin leak that nobody attributes to data at all.

Gartner's survey figures give a sense of the combined upside when this work lands: respondents to the survey behind its July 2024 analysis reported a 15.8% revenue increase, 15.2% cost savings and 22.6% productivity improvement on average. Those are averages across organisations that got somewhere, reported by the organisations themselves, and they sit in the same press release as the prediction that 30% of projects would be abandoned. Both halves are the same story: the work pays when it is sequenced, and it is abandoned when it is not.

The practical order we use is simple. Pick the process with the most hours in it. Fix only the records that process depends on, and count what that cleanup returns on its own. Build the automation. Measure against the baseline you recorded first. Then take the next process, which is now cheaper because some of its records were fixed by the last one.

That sequencing is the whole point of the strategy stage rather than an afterthought. Our strategy work decides which process is first and what has to be true before it can run, our engineering team builds what the plan calls for, and we run the result on our infrastructure from there.

Frequently asked questions

How much does it cost to get our data ready for AI?

Cost depends entirely on which process you are readying it for, which is why a general answer is worthless. Deduplicating one supplier list is usually days of work. Reconciling customer records across four systems that disagree is a project. Building a warehouse because somebody said you need one is an expensive answer to a question most businesses have not asked yet. BLACKSIG scopes the cleanup as part of the process it belongs to, on a call, rather than selling a data program as an entry fee. The useful number to establish first is what the manual process costs you per year, because that decides how much cleanup is worth doing.

How do you actually fix duplicate and inconsistent records?

Decide which system is the record for each kind of entity, then reconcile the others to it. For each duplicate cluster, pick the winning record by a rule rather than by hand, most recently updated or most complete, and keep the merged history where it matters. Then close the door: whatever created the duplicates, usually two people entering customers in two places, has to change or they come back. That last step is the one that gets skipped, and it is why some businesses clean the same data twice.

Is it better to fix our data first or automate first?

Fix the records the first automation touches, then automate, then repeat. Fixing everything first delays any return by months and usually stalls. Automating on top of records that are wrong produces confident wrong outputs at speed, which is worse than the manual process it replaced. The middle path is per process cleanup, and it works because each build pays for the tidying it needed.

Can we automate anything if our data is genuinely a mess?

Yes, and the place to start is a process whose inputs come from outside rather than from your own records. Reading incoming invoices, answering and booking calls, triaging inbound enquiries, extracting terms from contracts: these depend mostly on the document or the call in front of them, plus a small number of internal records. That makes them the usual first build for businesses whose internal data needs work, and the cleanup then happens process by process instead of as a project nobody can see the end of.

Do we need a data warehouse or a data lake before using AI?

Not for process automation. Warehouses earn their place when you need to analyse across systems, report on the whole business, or train something on history. Automating a process needs the records that process touches, reachable and correct. If a firm's first recommendation is a platform before it has looked at your processes, ask which process the platform is for and what it costs you today.

What about confidentiality if our data is sensitive?

Confidentiality is a separate question from data quality and it has its own answer. Where a system can send your data, who can see it, what is retained and what is logged are all design decisions made before a build starts, not settings adjusted afterwards.

Related resources

Find out where AI belongs in your business

We map how your business actually runs, decide where AI is worth using, then our engineering team builds what the plan calls for and we run it from there. We own the outcome, not the deliverable.