Your AI Is Only as Good as the Data Underneath It. Here Is How to Tell What Shape Yours Is In.
Updated: 3 days ago
An AI feature in a CRM does one thing: it reads the records as they are and produces something from them. A summary, a score, a next step, a forecast, a draft email. It does not know which fields were filled in carefully and which were filled in to close the screen. It does not know that half the deals in the pipeline closed months ago and nobody moved them. It reads what is there, and it is confident.
So the question before turning on any AI feature is not whether the model is good. It is what shape the data is in. We wrote about the general version of this in why your AI is not underperforming, your business context is. This post is the practical half: six things to measure in a CRM, with a threshold for each, and what the number tells you to fix.
Why the data, and not the model, decides.
A lead score built on a CRM where the industry field is empty on sixty percent of leads will score on the forty percent and generalise to the rest. A deal summary built on notes that were last written three months ago will summarise a deal that no longer exists. A next-best-action built on activity history where half the calls were never logged will recommend calling people who were called yesterday. None of these is the model failing. Each is the model doing exactly what it was asked with what it was given.

The companies whose AI rollouts disappoint almost always have a data problem that predates the rollout by years and that nobody measured, because until the AI feature arrived nothing read every record. People read the records they needed. The model reads all of them.
Six things to measure.
Every one of these can be measured from a report or a query in an afternoon. Do all six before you switch anything on.
One: completeness of the fields the feature will read.
For each field an AI feature will use, what fraction of records has it filled? Not the whole record; the specific fields. A lead scoring feature reads industry, size, source, title and activity. Measure those five. Below eighty percent on any of them, the feature will be scoring from a sample and pretending otherwise. The fix is rarely a data cleanup; it is finding out why the field is empty, which is usually that nobody needs it to do their job, and either making it required at the point where it is known or taking it out of the model.
Two: staleness of the records the feature will act on.
For open deals, open leads and active accounts, when was each last touched by a person? Count the ones untouched in ninety days. Above twenty percent of open deals, the pipeline is fiction, and any forecast or summary built on it inherits the fiction. The fix is a process: a rule that closes or flags anything untouched past a threshold, and a manager who looks at the flagged list weekly.
Three: duplicates.
How many contacts, accounts and leads exist twice? Measure by email, by domain, by phone, and by a loose name match. Above five percent, every count the model produces is wrong, every score is split across two records and every summary misses half the history. Zoho and most CRMs have merge tools; the fix is running them, then finding the entry path that created the duplicates, which is usually an import or a web form with no match rule.
Four: consistency of the values.
Pick five picklist and text fields the feature will read and count the distinct values. An industry field with four hundred distinct values, or a stage field that has been renamed three times so that old deals carry old names, is a field the model cannot group on. The threshold is whether a person could group the values by hand in an hour; if not, the model will not either. The fix is a controlled list, a one-time mapping of the old values, and a rule that stops free text in that field. We have written about the practices that keep a CRM clean and used, and consistency is the one that pays most for AI.
Five: activity coverage.
Of the interactions that happened last month, calls, emails, meetings, how many are in the CRM? Sample twenty deals and ask the owners. Below seventy percent, the activity history the model reads is a minority report, and any feature that reasons about engagement will be reasoning about the people who log things. The fix is integration, not discipline: email and calendar sync, telephony logging, so that the record is created by the event and not by the memory of it.
Six: ownership and permission shape.
Who owns the records, and can the feature see them? A model that runs under a user who cannot see half the pipeline will summarise half the pipeline. A model that can see everything, including fields the company considers restricted, is a governance problem before it is a data one. Measure what the feature's user can reach, and check it against what it should. The audit log is where you find out what it actually read.
What the numbers tell you.
Put the six on one page. Most CRMs we assess fail two or three of them, and the ones they fail are not random. Completeness and consistency fail together, because both come from fields that nobody needs to do their job. Staleness and activity coverage fail together, because both come from a process where the CRM is updated after the work rather than by it. Duplicates and ownership fail together, because both come from imports and integrations nobody owns.
That pattern is the point. The fix for a data quality problem is almost never a data cleanup, which restores the numbers for a quarter until the process that produced them runs again. The fix is the process: required fields at the point where the fact is known, records created by events rather than by memory, one owner for every entry path. A cleanup is worth doing once, on the day the process changes, so that the model starts from a clean floor.
Then turn the feature on, in this order.
Start with the feature that reads the fewest fields and the freshest records, which is usually a summary or a draft, because it fails visibly and cheaply. Move to scoring and next-step features when the completeness and consistency numbers are above threshold, because they fail quietly. Leave forecasting and autonomous actions until staleness and activity coverage are fixed, because they fail expensively. A company that turns them on in the other order gets a forecast it cannot trust in month one and a decision not to trust AI in month two.
How we assess it.
We run the six measurements as part of a readiness assessment, a few days of work that ends with a written score, the process fixes behind each failing number and the order to turn features on. If the data is not ready, the assessment says so and says what to fix first, and that is a better outcome than a rollout that gets blamed on the model. Where the fixes involve building anything, they come with a guaranteed estimate. The migration checklist covers the special case where the data is about to move, which is the one moment when a cleanup and a process change can happen together.
The short version.
An AI feature reads the CRM as it is. Measure completeness, staleness, duplicates, consistency, activity coverage and permission shape before turning anything on. Fix the process that produced each failing number, clean up once on the day the process changes, and turn features on in the order that fails cheaply first. The model is not the variable. The data is.
Find out what one connected Zoho system would change in the way you run.
In a no-risk discovery we look at your CRM, finance and operations systems and who owns each part, and show what a connected system would do differently. You pay only if you proceed. Or see how we approach it.
More on the same problem:
















Comments