Case study
How millions of spare parts got from the CRM onto the website, and why today we would start earlier at the data source.
An independent dealer in original industrial spare parts, around 50 employees. The product range sits in Salesforce, millions of parts from 2,226 manufacturers, and the website in four languages is the most important sales channel. It had to be built from this data, even though the data was not clean and Salesforce limits queries per month.

Entry 1
Starting point
This client marks the lower end of our target group. Around 50 employees is small for an AI project. But its dataset is larger than many a corporation's, and the business depends on it. Anyone looking for a spare part for a plant searches by part number and manufacturer, finds a product page and sends an enquiry. The more parts can be found, the more enquiries come in.
The product range is maintained in Salesforce, by sales, for quotations, not for a website. And it showed in the data: placeholders, internal IDs, drawing numbers in the name field, discontinued products that were never deleted. Unfiltered, all of this would have reached the website.
Apart from the company name, the manufacturer pages had identical text, the same two-line blurb 2,226 times. Such pages are worthless to search engines, and to customers too. And Salesforce limits queries per month: search engine crawlers visiting millions of pages used up the allowance before a customer saw the site. An enquiry with twenty lines produced twenty separate emails that someone in the sales office had to piece together.
Entry 2
Figures before
| Manufacturers in the range | 2,226 |
|---|---|
| Parts | millions, maintained in Salesforce |
| Website languages | 4 |
| Manufacturer texts | 2,226 identical two-line blurbs |
| Data quality | placeholders, internal IDs, drawing numbers in the name field, discontinued products |
| Query allowance | used up by crawlers before customers arrived |
| Enquiry with 20 lines | 20 separate emails |
Entry 3
The path
About ten weeks to the first weekly rollout, automatic operation after that. The week figures are rounded.
- Weeks 1 to 2
Counting before building
We analysed the stock: which kinds of data junk exist and how often, and where the Salesforce queries come from. Result: the biggest consumer was not customers but crawlers.
- Weeks 3 to 4
Filters before processing
A filter recognises placeholders, drawing numbers in the name field and discontinued products and keeps them out of everything built from the data: pages, sitemaps, texts. The data in Salesforce stays unchanged, and sales keeps working as usual.
- Weeks 3 to 6
Caching and throttling
Caching in the database instead of in files, lifetimes per data type, negative caching for dead URLs, a search budget per hour and crawler rules at the network edge. The site did not get any slower as a result.
- Weeks 5 to 8
Manufacturer texts anchored in facts
Each manufacturer gets its own text in German and English: introduction, company profile, spare parts notes, meta description. The fact anchor is the real product range on the live page, and the positioning as an independent dealer appears in every text. Quality checks run before go-live; if there is an error, a bit-identical rollback follows.
- Week 9
Trial run
The first ten manufacturers go through the whole workflow and are also proofread by hand, by us and by the client.
- From week 10
Weekly rollout with no manual work
Ten manufacturers a week, at night, sorted by search demand, down to rank 300, because the top 300 account for 86 percent of manufacturer traffic. A weekly report arrives on Fridays. A stop switch halts the rollout without breaking anything.
Entry 4
Result
Today, the Salesforce data produces 3.3 million product pages in four languages, without placeholders or discontinued products. Salesforce queries have fallen by 71 percent, and the allowance holds. The manufacturer texts go live in the automated weekly rollout, each with its own content instead of a two-line blurb. An enquiry with twenty lines arrives as one email, with a Salesforce ID per line and a data block that the next stage of automation can read directly.
| Metric | before | after |
|---|---|---|
| Product pages | with data junk | 3.3 million, filtered, four languages |
| Manufacturer texts | 2,226 identical two-line blurbs | a text of its own per manufacturer, ten a week |
| Salesforce queries | allowance used up | minus 71 percent |
| Checks before go-live | none | eleven quality checks, bit-identical rollback |
| Enquiry with 20 lines | 20 emails | one machine-readable email |
| Monitoring | none | weekly report on Fridays, at most one error message per hour per error type |
Operating mode: product data contains no personal data. Text research runs via cloud models under contract, rollout and checks run on our infrastructure in Germany, and enquiry data stays with the client. Approval works through the weekly report: the client reads it and can stop the rollout at any time.
Entry 5
What went wrong
- Outdated group affiliations. For some manufacturers, the first text versions named parent companies that were no longer correct after acquisitions or sales. For an independent dealer, this is doubly sensitive, because a wrong affiliation can look like a distribution partnership. The lesson: the research rule now says to name a group affiliation only with a date and a source, otherwise it is left out.
- A missing check in the first week. In the first rollout, the workflow did not yet check whether umlauts were encoded correctly. Three texts with character errors got through. We corrected them and added the check as the eleventh. The lesson: character encoding is not a minor detail; it is the first check in every multilingual workflow.
Entry 6
What we would do differently today
We would start with the checks and then have the texts written, not the other way round. The ten checks at the start were derived from the errors we expected. The eleventh came from an error we did not expect. Today, before the first text, we write a list of everything that can be wrong with a text and have each item checked.
For facts about third parties, such as group structures, headquarters or founding years, one rule applies from the outset: source and date, or not at all. That costs a few sentences of content and saves corrections.
And we would talk to sales earlier about data maintenance in Salesforce. The filter keeps the junk off the website, but it keeps being created at the source. A few mandatory fields and picklists when records are created would have taken half the work off the filter.
Entry 7
Key facts
| Industry | Independent trade in original industrial spare parts |
|---|---|
| Employees | around 50 |
| Systems | Salesforce, WordPress, Cloudflare, Google Search Console, Postmark |
| Services | Putting data in order, Texts and content at scale, ongoing operation |
| Operating mode | Product data without personal data; text research via cloud models under contract; operation on our infrastructure in Germany |
| Duration | about ten weeks to the first weekly rollout, around 30 weeks of rollout down to rank 300, in ongoing operation ever since |
Read more
Further reading
Handover
The first step is a 30-minute call.
You tell us about the workflow that costs you the most time. We tell you honestly whether AI pays off there and what the next step would be. Whether a workflow analysis follows is up to you.