Why volume changes the problem
A text that a person writes with AI and then reads is a text. 4,000 texts that nobody reads individually any more are a workflow, and a workflow needs different safeguards. The arithmetic is simple: if one text in a hundred contains an error, that means 40 wrong pages across 4,000 products, and nobody knows which. With technical products, a wrong value is not a matter of style but a complaint or a liability issue.
So the question is not whether a language model can write good product texts. It can. The question is how to make sure that no bad one goes online without reading every single one. The answer has three parts: a fact anchor, a rulebook and automatic checks. How we build such workflows is described under Texts and content at scale; here we deal with the principle, which you can demand from any provider.
The fact anchor: no sentence without a source
Language models fill in what sounds plausible. For product texts, that is their most dangerous property. The fact anchor reverses it: the text may only contain statements that appear in a defined source, such as the PIM field, the manufacturer's data sheet or the product range on the live site. What the source does not provide is not invented but reported as a gap.
This has a pleasant side effect: the list of gaps is a work list for product management. After the first run you know which products lack weight, protection rating or voltage, as a complete list rather than a random selection.
At an independent dealer in original industrial spare parts, the fact anchor for the manufacturer texts is the actual product range on the live site. Group affiliations are researched, not guessed. And the data is filtered before processing: placeholders, drawing numbers in the name field and discontinued products are thrown out before any text is produced from them. A fact anchor is only as good as the source it is tied to.
The rulebook: what a text must, may and must never do
Before the first text is produced, the rules are written down. The rulebook is a document that your editorial team can read and change, not a hidden prompt. It typically contains:
- Structure per text type: short description, long text, benefits, meta description, each with a length.
- Fixed terms from the glossary, per language, and terms that must never appear.
- Prohibited statements: superlatives without evidence, environmental claims without proof, promises about delivery time or warranty.
- Mandatory statements that appear in every text, such as a note on independence from the manufacturer.
- Tone of voice in two or three sentences, with a good and a bad example.
Environmental claims deserve particular attention. EU Directive 2024/825 (EmpCo) only permits generic claims such as "environmentally friendly" or "climate neutral" with recognised proof; member states have applied the requirements since 27 September 2026. If such a statement is not covered in the rulebook, it will be in thousands of texts after the first run. How to check an existing text inventory for this is described in the article The EmpCo Directive. This article is not legal advice.
The checklist: eleven checks before every go-live
At this spare parts dealer, every manufacturer text goes through eleven quality checks before it goes live. Texts are rejected for, among other things, wrong characters, missing umlauts, overly long meta texts, invented figures, a missing independence notice and duplicate manufacturers. After insertion the file is checked in full; if anything is wrong, there is a bit-identical rollback and no go-live. Generalised to product texts of all kinds, this gives the following checklist. Every check runs automatically, and each has a clear outcome: passed, or rejected with a reason.
| No. | Check | Rejected if |
|---|---|---|
| 1 | Character set | control characters, broken encoding or foreign script characters appear in the text |
| 2 | Umlauts and special characters | "ae" appears instead of "ä", or accents of the target language are missing |
| 3 | Lengths | title, short text or meta description exceed or fall short of the set length |
| 4 | Figures against the source | a figure, measured value or metric does not appear in the source |
| 5 | Units | a unit does not fit the characteristic or has been converted although the rulebook does not allow it |
| 6 | Prohibited terms and statements | a term from the block list or an unsubstantiated environmental or superlative claim appears |
| 7 | Mandatory statements | a required notice is missing, for example on independence or safety requirements |
| 8 | Glossary | a product term in the target language differs from the one set in the glossary |
| 9 | Language | the text is wholly or partly in the wrong language |
| 10 | Duplication | two products or manufacturers receive the same or almost the same text |
| 11 | Full check after import | the page or file is not complete and valid after insertion; in that case, rollback to the previous state |
Checks 4 and 10 catch the most serious errors. Check 4 prevents invented values. Check 10 prevents 2,000 products from turning into 2,000 variants of the same text. How important that is can be seen from the starting point at the spare parts dealer: 2,226 manufacturer pages had identical text apart from the company name, the same two lines 2,226 times.
One lesson from our text projects: the checklist is never complete on the first day. Every error that still slips through is therefore turned into a new check the same day, so that it cannot get through a second time. A checklist that looks exactly the same after three months as it did at the start is not being maintained.
Trial run, waves, stop switch
Even with a complete checklist, no text inventory goes online in one step. This sequence has proved its worth:
- A trial run with around 100 texts that your editorial team reads in full. Whatever does not fit becomes a rule or a check. Only when the sample passes is the full inventory processed.
- Rollout in waves, sorted by importance: first the products or manufacturers with the most search demand or the most revenue.
- Samples per wave instead of reading every text. The sample gets smaller if the error rate stays low, and larger again if it rises.
- A report per wave: what was produced, what was rejected and why, which gaps have been reported.
- A stop switch that halts the rollout without destroying anything, and a rollback to the previous state.
At the spare parts dealer, ten manufacturer pages go live each week, overnight and without manual work, sorted by search demand down to rank 300, because that is where 86 percent of manufacturer traffic lies. The weekly report arrives on Fridays. The whole story is in the case study on product data from Salesforce.
Multiple languages: the glossary decides
For several languages there are two routes: produce each language directly from the source data, or translate a checked source text. For technical products the second route is usually the safer one, because the fact check only has to run once against the source, and the translation then only has to be checked against the approved text. In both cases the specialist glossary is decisive: a product term must have the same name in every language, on every page.
Two rules save trouble. First: only empty fields are filled; editorially checked texts remain untouched. Second: before the first run, the glossary is checked for double assignments, meaning terms that denote two different things in one language. This is how a fuel cell manufacturer maintains error codes, help texts and app content centrally and has them AI-translated into eleven languages, using the company's specialist glossary (Case study: regulated texts across ten domains).
Who approves, and what has to be labelled
Approval stays with you. The only open question is in what form. For product texts, after a successful trial run, a sample per wave is usually enough. For texts with legal weight, such as reworded environmental claims, we recommend approval of each passage by a person. The log records which text was produced according to which rule and who approved it.
On labelling: Article 50(4) of the EU AI Act, Regulation (EU) 2024/1689, requires disclosure of AI-generated texts that are published to inform the public on matters of public interest. This does not apply if people have reviewed the text and a person bears editorial responsibility. In our assessment, product descriptions in a shop are generally not covered. Clarify the question for your case with your legal advisers. This article is not legal advice.
What this means for you
Before you talk about models, clarify three things: which source is authoritative for each product characteristic, who helps write the rulebook, and who approves. Hold the eleven checks up against your existing texts and strike out what does not apply. What remains is the benchmark for every provider, including us.
Your next step: take the potential check. It shows you immediately on the page whether a workflow pays off for your text inventory.
Further reading
- The EmpCo Directive: how to find environmental claims across your entire text inventory and reword them so they can be substantiated
- Before AI is let into the CRM: sorting out duplicates, mandatory fields and spellings
- Service: texts and content at scale
- Case study: product data from Salesforce
- Industry: AI in distribution, with data sheets into the PIM
