Why AI does not repair bad data but spreads it
Anyone who has worked in sales for ten years knows that "Mueller GmbH" and "Müller Antriebstechnik GmbH & Co. KG" are the same company. They skip over the test entry "XXX please delete" and know which of three contacts is still responsible. This knowledge is not written down anywhere. A workflow that automatically assigns enquiries, selects campaigns or prepares quotations does not have it. It takes the data literally.
This leads to a simple rule: an automatic workflow does not make an error in the data smaller; it just uses it more often. A duplicate that bothers a person once a quarter bothers a workflow that assigns 300 enquiries overnight, every night. That is why clean-up comes before every AI project that works on customer data, not after it. For us it is a separate service with its own fixed price: Putting data in order.
The five types of error found in almost every CRM
The names change from company to company, the patterns do not. When we see a dataset for the first time, we look for these five.
- Duplicates. The same company or person several times, created by different colleagues, via forms, trade fairs and imports from the ERP. The spelling differs, the address is an old one, the email a different one.
- Junk data. Test entries, placeholders such as "unknown" or "n/a", aborted imports with half-filled rows, internal notes in the name field.
- Empty mandatory fields. Industry, country, company size, responsibility: fields that sales filters by are empty for some records, because they were allowed to be empty when created.
- Inconsistent values. "Deutschland", "DE", "Germany" and "BRD" in the same field. Salutations in four variants. Industries as free text instead of from a list.
- Outdated data. Contacts who have left the company, companies that have been taken over or renamed, addresses from before the move.
Count first, then clean
Before a rule is written, you need the findings: the data in figures. They answer whether a clean-up pays off and where it should start. Most CRM systems provide the figures through an export and an evaluation that your IT can produce in a day. Five questions are enough:
- How many records are there per object (companies, contacts, leads), and how many of them have been touched in the last 24 months?
- How many possible duplicates does a simple comparison of normalised company names and email domains find?
- How complete are the fields that sales and marketing filter by?
- How many different values are there in fields that should really have a fixed list, such as country, industry, salutation?
- How many email addresses are demonstrably undeliverable, and how many contacts are no longer linked to an active company?
The order follows from these five figures. A dataset with few duplicates but an empty industry field needs enrichment, not merging. A dataset with 40 spellings for Germany needs standardising first, because otherwise the duplicate comparison misses matches.
Check rules for duplicates: what is merged automatically and what is not
The core problem of any duplicate search is not finding but deciding. Too strict, and half the duplicates remain. Too loose, and two different companies with similar names become one, sales history and all. The way out is graded rules with three outcomes: merge, put before a person, keep separate. The following table is a starting point that you can adapt to your data.
| Characteristic | What it means | Outcome |
|---|---|---|
| Email address identical (lower case, no spaces) | the same person | merge automatically |
| VAT ID or commercial register number identical | the same legal entity | merge automatically |
| Normalised company name and postcode the same | name without legal form, umlauts resolved, special characters removed; very probably the same company | merge automatically if no other field contradicts it |
| Normalised name the same, town different | may be a branch or subsidiary | put before a person |
| Email domain the same, company name different | group, brand or change of name | put before a person, never automatically |
| Name similar, no other match | typo or abbreviation, such as "Hofmann" and "Hoffmann" | put before a person |
| Free email domain as the only shared characteristic | says nothing about the company | keep separate |
| One record has open sales opportunities or contracts | history comes before tidiness | put before a person; the responsible salesperson merges |
Two further rules belong here that have nothing to do with similarity. First: when merging, it is not the older or the newer record that wins, but the more reliable value for each field. You set the order of precedence, for example address from the ERP before address from the web form. Second: every merge is logged, with both original records, so that it can be undone.
Spellings and mandatory fields: deriving rules from the data
The same principle applies to standardising as to duplicates: one value per meaning, and the mapping is in a table that you read and confirm. For countries this is quickly done. For industries it takes longer, because the data often contains its own terms that sales genuinely uses. Those belong in the target list, not in the bin.
For empty mandatory fields there are three routes, in this order. First, derive from existing data (country from postcode and dialling code, industry from other contacts at the same company). Then enrich from public sources (the company's website, the commercial register). And where neither is enough, leave the field empty and mark it. An invented value is worse than an empty field, because nobody distrusts it any more. Every enriched value therefore gets a source reference.
Where a language model helps and where rules are enough
A large part of the clean-up needs no AI. Normalising, comparing, counting and merging according to fixed rules is done by ordinary software, quickly, traceably and entirely within your system. That is also the simpler route for data protection. A language model pays off where rules fail because of language:
- Reading free-text fields in which notes, telephone numbers and contacts are all mixed up.
- Assigning the industry from a company website's self-description to a fixed list.
- Preparing an assessment with reasons for the borderline cases referred to a person, so that the person decides faster.
- Recognising junk data that follows no pattern, such as a drawing number in the name field.
Customer data is personal data. Where a model reads it, we use EU data centres or open models on your own servers, whichever you decide. What this looks like for each service is described under Data sovereignty.
The run: backup, trial run, log
Once the rules are in place, the actual run is the smaller part of the work. It always follows the same order:
- A complete backup of the data before anything is changed.
- A trial run on a sample of a few hundred records. Your team checks the result, and deviations become new rules.
- A run over the entire dataset, with a log per change: old value, new value, rule applied, time.
- The borderline cases put before the people responsible, bundled in a list, not as individual emails.
- A check using the same five figures as in the findings, before and after side by side.
One point is often overlooked with cloud CRMs: Salesforce, HubSpot and others limit the number of interface requests per period. A run over hundreds of thousands of records must be built so that it does not use up the allowance, otherwise the form on the website stops working the next morning. What this looks like with millions of records is shown in the case study of a spare parts dealer that draws its product data from Salesforce: there the junk data is thrown out before any further processing, and the requests are built so that the limit holds.
Keeping the data clean
Clean data soon degrades again if only the existing records are cleaned and not the incoming data. Three measures keep it clean:
- The same rules apply to every import and every form: country code instead of free text, a search for the existing company before every new record.
- Only the fields that are really needed are mandatory. A mandatory field that nobody can fill sensibly produces placeholders.
- A recurring check counts new duplicates and empty fields every month and reports when the figures rise.
Clean data is also the prerequisite for the most common AI workflow in sales, the automatic assignment of incoming enquiries. If the workflow finds a company three times in the CRM, with three different people responsible, it cannot distribute correctly. How this workflow is built is described in the article Lead qualification with AI in B2B sales.
There is also a legal side. Article 5(1)(d) GDPR requires personal data to be accurate and, where necessary, kept up to date; point (e) requires that it be kept no longer than necessary for the purpose. A clean-up that finds outdated contacts therefore also supplies the list for the deletion review. Whether and when you have to delete is something to clarify with your data protection officer. This article is not legal advice.
What this means for you
If you are planning an AI workflow on your customer data, count first. The five figures from the findings show whether your data is good enough or whether the clean-up becomes the first workflow. You can apply the check rules for duplicates to a sample with your IT even without us; afterwards you will know how much work there is in your data.
Your next step: take the potential check. It also asks about the quality of your data and shows you the result immediately on the page.
Further reading
- Reading, classifying and distributing enquiries automatically: lead qualification with AI in B2B sales
- AI potential analysis: which workflows pay off with AI, and which do not
- Service: putting data in order
- Case study: millions of product records from the CRM
- Industry: AI in distribution, where master data built up over many years is the norm
