CRM Data Hygiene: The Maintenance Job Nobody Is Assigned.
Ask a mid-market company when it last cleaned its CRM and the answer is a date: eighteen months ago, before the last system change, when the new sales lead started. Ask who keeps it clean and the answer is a pause. The cleanup was a project. Keeping it clean would be a job, and the job was never created.
That gap is the whole problem. Data decays continuously, at a rate you can estimate, and cleanup happens episodically, when the pain gets loud enough. Between the two the data is always getting worse, and now there is an AI or an automation sitting on top of it, built against the clean snapshot, degrading from the day it went live.
How fast the data decays.
The decay has three sources, and they do not stop.

People move. In a B2B contact base, a meaningful share of contacts changes job, title, email or employer every year. The record does not know. It keeps the old title, the old email bounces, and the next campaign reports a delivery rate that is really a decay rate.
Companies change. They rename, merge, get acquired, move. The account record in the sales system keeps the old name while finance types the new one from the purchase order, and now the customer exists twice, which is the pattern that breaks every cross-system report.
Duplicates accrete. A web form creates a second lead for a known contact because the email was typed differently. An import brings in three hundred accounts, forty of which already existed under slightly different names. A rep creates a contact rather than searching. Each event is small, and there are thousands of them a year.
None of this is a failure of the people. It is what a record does when the world it describes keeps moving. The question is only whether anyone is watching.
Why the cleanup project does not fix it.
A cleanup project does three weeks of work and produces a clean database. It is satisfying and it is the wrong shape, because the day it finishes the decay resumes at the same rate, with nobody watching, and in eighteen months the company runs another one. We described how to measure the state of the data in another piece. Measuring it once tells you where you are. It does not keep you there.
The cleanup project has a second cost that is easy to miss. It usually happens right before something new is built on the data: a reporting layer, an automation, an agent. So the new thing is designed and tested against the cleanest the data will ever be, and then meets reality in month three, when the decay has reintroduced the duplicates and the agent chases the wrong contact. The project made the degradation harder to see, not easier.
The job, as a job.
Hygiene is a maintenance job, and a maintenance job has three parts: an owner, a cadence and a checklist. None of them is technical.
The owner is one named person, usually in sales operations or the CRM administrator, with the responsibility written into their role. Not a committee, not the vendor, not everyone. The test is that when a duplicate appears in a report, there is a person whose Monday it ruins.
The cadence is an hour a week. Not a quarter-end push. An hour, on the calendar, every week, in which the owner works the four checks below. The hour is enough because the decay per week is small; it only becomes a project when it compounds for a year.
The checklist is four standing reports, built once in the CRM and read weekly.
Likely duplicates. Leads and contacts matching on email or on name plus company, created in the last seven days, against the existing base. The owner merges them, by the merge rule the company wrote down, and notes which door they came through.
Bounced and stale. Contacts whose last email bounced, and contacts with no activity in twelve months whose company has changed in the account record. The owner verifies or retires them.
Required fields empty. Records created this week missing the fields the company said were required: industry, owner, source, country. Each one is a door that let a bad record in, and the owner fixes the door, not just the record.
Accounts without an id in the other system. For a company with a connected accounting or help desk system, the accounts created this week that do not yet carry the shared customer id. This is the check that keeps the cross-system reports honest, and it is the one nobody thinks of as hygiene.
Four reports, an hour, a name. That is the job.
What the job makes possible.
The reason to create it is not tidiness. It is that everything the company wants to build next depends on it holding.
A report that joins the sales system to the ledger only stays right if the customer id check runs every week. A lead-routing automation only routes correctly if the required fields are filled at the door. An agent that follows up leads only follows up the right person if the duplicates are merged before it reads them. The same company that will spend five figures on the agent will not create a job for an hour a week, and then will blame the agent when it acts on the duplicate.
There is also a measurable return that has nothing to do with AI. Campaign delivery rates stop drifting down. The forecast stops counting the same deal twice. The month-end reconciliation between systems stops finding customers that exist under two names. These are the complaints that trigger the next cleanup project, and the job makes the project unnecessary.
Where the dedupe rule comes from.
The job needs one input the owner cannot invent: the merge rule. When two records match, which one wins, which fields are kept from the loser, and what happens to the activity history. That is a policy decision, made once by someone with the authority to decide that the sales system's name beats the accounting system's name, or the reverse. Without it the owner either merges inconsistently or does not merge at all, and the weekly hour produces a list and no action.
Write the rule in a page. Configure it as the CRM's duplicate check so new records are caught at creation, which is cheaper than merging later. Then the weekly report is only the ones that got past the door.
Starting it.
The first week is longer than an hour, because the backlog has to be cleared once, and that is the only time a cleanup project is the right shape: as the start of the job, not instead of it. After that the four reports take their hour, and the owner's name on the calendar is what keeps the data in the state the AI was built on.
We set up the four reports and the merge rule, and we document the job so the owner can run it without us. Discovery is no-risk: we read the current state of the data, show you the decay rate, and you pay only if you go ahead.
Know whether your data and processes are ready for AI before you pay for it.
The AI Readiness Review is a 90-minute working session plus a written scorecard across data hygiene, documented process, permissions and ownership. Fixed scope, no obligation. Or see how we approach it.
More on the same problem:
















Comments