Zweite Schicht DE, Deutsche Fassung Book a call

4 Oct 2026

Reading time10 minutes

AI data protection in companies: where models should run for which data, on your own servers, in Europe or in the cloud

At the project meeting the data protection officer says one clear sentence: customer data does not go into the cloud. Two floors down, marketing has been copying product texts into a freely accessible chat window for months, and in sales the occasional customer email ends up there too. Both show that the question is being asked the wrong way: it is not about AI on your own servers or in the cloud, but about which data may go where.

, reading time 10 minutes, by Zweite Schicht

Sketch of three stations on one pipe: a server cabinet with a key in its lock, a data centre building with cooling units on the roof and a contract folder with a seal.

Why "local or cloud" is the wrong question

Anyone who talks about AI and data protection quickly ends up with two camps. One wants to run everything on its own servers, so that nothing gets out. The other wants the strongest models, and those are run by the big providers. Both positions have a core of truth, and both are misleading if they are meant to apply to all workflows.

A company processes very different data. Product descriptions that are on the website anyway need a different level of protection from job applications, employment contracts or the commercial terms of a major customer. Anyone who treats everything the same either protects the harmless data too much and gives up quality, or protects the sensitive data too little. It makes more sense to decide for each workflow which data is involved and choose the operating mode accordingly.

We distinguish three operating modes. Each has its place, and in most companies several end up running side by side.

The three operating modes at a glance

1. On your own servers

Open models such as Llama or Mistral run on your own hardware or on a dedicated server in Germany that only you and your service provider can reach. No data leaves the network.

  • Strength: full control over data and model, no fees per request, no dependence on a provider's price list.
  • Limit: the open models are good for language at volume, such as classifying, summarising and extracting fields. For difficult individual cases and demanding texts they are usually inferior to the large cloud models.
  • Effort: set-up, computing power with sufficient GPU memory, maintenance and updates. The larger the model, the more hardware it needs.

2. EU data centre

Large models through providers with processing in the EU, such as Google Vertex AI in Frankfurt or Belgium, Microsoft Azure in Germany or Mistral in France. Covered by a data processing agreement, with no use of your data for training and storage only for the duration of processing.

  • Strength: strong models with processing inside the EU, quick to set up, no hardware of your own.
  • Limit: with providers headquartered in the USA, the question remains whether US authorities can demand data on the basis of the CLOUD Act of 2018, even if it is stored in Europe. Many data protection officers weigh up this residual risk and consider it acceptable for customer data in normal cases, but not for particularly sensitive data.
  • Effort: contract review, configuring the region, fees per token.

3. Cloud under contract

Models from Anthropic, OpenAI or Google directly through the provider's interface, with a data processing agreement and a contractual undertaking that inputs are not used for training. This is not the same as the freely accessible chat window into which employees copy texts.

  • Strength: the best results for language, the newest models, very little set-up effort.
  • Limit: processing may take place outside the EU. Suitable where no personal data flows, or where you approve it in writing after weighing up the risks.
  • Effort: low. Fees per token, contract review, documentation of the third-country transfer if personal data is involved.

Which data belongs in which operating mode

The following table is our starting point in every analysis. It does not replace weighing up the individual case, but in most cases it answers the first question. Go through your planned workflows row by row and note which types of data are involved. The strictest row determines the operating mode, unless the sensitive parts can be removed beforehand.

Which data belongs in which operating mode
Type of dataRecommended operating modeReason
Product texts, data sheets, website contentCloud under contractNo personal data, usually public anyway. What counts here is text quality.
Translations with a specialist glossaryCloud under contractAs with product texts, as long as no customer names or commercial terms appear in the text.
Customer enquiries, emails, CRM recordsEU data centrePersonal data in normal business dealings. Processing in the EU under a data processing agreement.
CRM master data during clean-upWithout a language model, otherwise on your own servers or in EuropeDuplicates and spellings can often be resolved with rules in the system. Where a model helps, it stays close to the data.
Job applications, personnel filesOn your own servers or in Europe after weighing upSensitive personal data, co-determination by the works council, and obligations under the EU AI Act, Regulation (EU) 2024/1689, for selection decisions.
Health data, data under Article 9 GDPROn your own serversSpecial categories of personal data, high requirements for any processing.
Prices, commercial terms, contractsOn your own servers, or replace beforehandTrade secrets. Under section 2 of the German Trade Secrets Act (GeschGehG), protection requires appropriate confidentiality measures.
Design data, formulationsOn your own serversThe core of the company's value. No quality advantage of the cloud justifies the risk here.
Knowledge base from your own documentsOn your own servers, queries depending on contentThe collection itself stays with you; only the excerpt needed for an answer goes to the model.

Our customer cases, almost all of them hybrids, show what this looks like in practice. At a technology distributor, the knowledge base is on its own infrastructure, the enquiries run in NetSuite and the chat runs via a cloud model under contract. At a spare parts dealer, product data without personal data is researched via cloud models, while enquiry data stays with the customer. At a fuel cell manufacturer, texts without personal data run via cloud models under contract, while personal data stays in Salesforce and SAP SuccessFactors.

Replacing beforehand: the underestimated intermediate step

Many workflows contain only a few sensitive passages. A product enquiry consists of a technical question and a signature with name, telephone number and company. A quotation text contains a description and a line with the discount. In such cases it is often better to replace the sensitive parts before processing than to move the whole workflow to a weaker operating mode.

Technically, this means: a pre-processing step on your own servers recognises names, addresses, customer numbers or prices and replaces them with placeholders. The model works with the cleaned text, and afterwards the workflow puts the original values back in. The GDPR calls this pseudonymisation in Article 4(5) and recommends it as a protective measure in Articles 25 and 32. Under Recital 26, pseudonymised data is still generally regarded as personal data; but the risk is considerably reduced, and for prices and commercial terms this step is often the simplest way to keep them on your own servers.

The legal background

The choice of operating mode has consequences for the documents your data protection officer needs. The most important provisions, as of October 2026:

  • GDPR Article 28: anyone who processes personal data on behalf of another needs a data processing agreement. This applies to your service provider and to every model provider involved as a sub-processor.
  • GDPR Articles 44 to 46: transfers to third countries need a legal basis, such as an adequacy decision under Article 45 or standard contractual clauses under Article 46. For certified US companies, the adequacy decision on the EU-US Data Privacy Framework has applied since July 2023. It has been challenged in court; whether it will survive in the long term remains open.
  • GDPR Article 9: health data and other special categories may only be processed under narrow conditions.
  • GDPR Article 35: where there is likely to be a high risk for data subjects, a data protection impact assessment is required. That can be the case for AI workflows involving personnel data.
  • EU AI Act: the operating mode does not change the obligations under the EU AI Act, such as the transparency obligations under Article 50. Which obligations arise for which workflow is described in the article The EU AI Act for mid-sized companies.

This overview is not legal advice. Have the classification of your workflows checked by your data protection officer or your legal advisers; the data flow sketch for each workflow is the basis for this.

Keeping models interchangeable

A second question is closely tied to the operating mode: what happens if a provider raises prices, retires a model or a better one appears? Workflows should be built so that the model is interchangeable. The rules, the checks and the log belong to the workflow, not to the model. If the model changes, the same checks run over a sample before the new model takes over.

This is also a safeguard for the operating mode itself. If the legal situation for third-country transfers changes, a workflow can be moved from the cloud to an EU data centre or onto your own servers without rebuilding it. The open models are getting better; what only works well in the cloud today may run on your own servers in a year's time.

Eight questions to ask any AI provider

These questions let you check whether a provider merely promises data sovereignty or can prove it:

  1. Which data flows to which model provider in which workflow, and can you get this as a diagram?
  2. In which region is the data processed, and where is that set down in the contract?
  3. Does the contract state that your data will not be used to train third-party models?
  4. Is there a data processing agreement under Article 28 GDPR before the first access to data?
  5. Which sub-processors are involved, and on what basis does any third-country transfer take place?
  6. How long are inputs stored, and how is deletion demonstrated?
  7. Can sensitive parts be replaced before processing?
  8. Can the model be changed or the operating mode moved without rebuilding the workflow?

Our answers to these questions, including the data flow for each service and the sentence that appears in every one of our contracts, can be found on the Data sovereignty page. For workflows with customer enquiries and knowledge bases, the Sales and communication page describes how the operating modes are combined there.

What this means for you

You do not have to choose between data protection and good results. If you record for each workflow which data is involved, the operating mode usually follows by itself: public content to the cloud under contract, customer data to an EU data centre, secrets and particularly sensitive data onto your own servers. Where only a few passages are sensitive, replace them beforehand. And the freely accessible chat window into which customer emails may be copied today should be replaced by a workflow with a contract.

Your next step: go through the checklist for introducing AI on a sound legal footing with your data protection officer. The data protection section covers data processing agreements, third-country transfers and deletion rules, and the technology section covers the choice of operating mode.

Further reading

Handover

The first step is a 30-minute call.

You tell us about the workflow that costs you the most time. We tell you honestly whether AI pays off there and what the next step would be. Whether a workflow analysis follows is up to you.

Book a callApproach and prices