Accessibility settings

Text size

100%

AI 4 min read

Data anonymization

Also known as: anonymization, PII anonymization, PII redaction, pseudonymization

Definition

Data anonymization is processing personal data so that no one can identify the person by any means reasonably likely to be used. Anonymous data falls outside the GDPR; pseudonymized data, which can be re-linked with separately kept information, remains personal data under the GDPR.

Cite this entry

Text

"Data anonymization". Order Group, Software glossary, 10 October 2026. https://ordergroup.co/glossary/data-anonymization/

HTML

<a href="https://ordergroup.co/glossary/data-anonymization/">Data anonymization</a> - Order Group

How data anonymization works

The GDPR separates two outcomes. Pseudonymization, defined in Article 4(5), is processing personal data so that it can no longer be attributed to a specific person without additional information, provided that information is kept separately and protected by technical and organizational measures. Pseudonymized data is still personal data. Anonymous data is information that does not relate to an identifiable person, and Recital 26 says the GDPR does not apply to it. Whether a person is identifiable depends on all the means reasonably likely to be used, singling out included, taking into account cost, time and available technology.

The Article 29 Working Party (the predecessor of the EDPB) tested anonymization techniques against three risks in its Opinion 05/2014: singling out one person in a dataset, linking two records about the same person, and inferring a fact about a person. It grouped techniques into randomization (noise addition, permutation, differential privacy) and generalization (aggregation, k-anonymity, l-diversity, t-closeness). Each technique leaves some risk, so the result is assessed per dataset and repeated when the data or available tools change.

Free text is harder than tables. Names, addresses, ID numbers and phone numbers are detected with named entity recognition models, regular expressions and checksums. Microsoft Presidio, an open-source toolkit, works this way: an analyzer finds entities and an anonymizer replaces, masks, hashes or encrypts them. Detection is probabilistic. It misses unusual spellings and inflected names, and it cannot see indirect identifiers: a job title, a small town and a rare diagnosis in one letter can point to one person even after the name is gone. Replacing names with placeholders is therefore usually pseudonymization, and only an assessment of the remaining text shows whether it is anonymous.

What data anonymization means for your software

In a product with a language model, personal data appears in three places: in the text sent to the model, in what the system stores, and in any corpus built for training or evaluation.

  • Before the model: if the model runs on your infrastructure (a self-hosted LLM), the text may go to it unchanged within the purpose and legal basis of the processing. If it goes to an external API, redact first, or pseudonymize and put the names back into the answer, with the mapping kept separately and encrypted.
  • Stored texts and history: prompts, outputs and uploaded files are personal data. The system has a retention period, a scheduled deletion job, and a test that proves the job deletes files as well as database rows.
  • Logs: the application, the inference server and the error tracker each log requests. Decide what they keep, mask text fields, and give logs a shorter retention than the content itself.
  • Training and evaluation corpus: only texts covered by a legal basis for that purpose (for health or other special-category data, usually explicit consent), anonymized before export, with a person checking a sample for anything the detector missed.
  • Detection quality is measured on your own documents and language: what share of the real names, addresses and numbers the detector caught. A demo on clean English sentences says little about scanned Polish letters.
  • Encryption at rest protects stored texts but does not anonymize them, because whoever holds the key can read them.
Anonymization, pseudonymization and encryption in a system with a language model
AspectAnonymizationPseudonymizationEncryption
Can the person be identified againNo, by any means reasonably likely to be usedYes, with the separately kept informationYes, with the key
GDPR appliesNo (Recital 26)Yes, for whoever holds the additional information (Article 4(5), Recital 26)Yes
Typical useTraining and evaluation corpus, statisticsSending text to an external model and restoring names in the answerStored documents and history
Main riskContext in free text still identifies the personThe mapping table leaksKey management and access

Rules and standards

Besides Article 4(5) and Recital 26, the GDPR touches anonymization in Article 5(1)(e), storage limitation, which is the reason for a retention period; Article 25, data protection by design and by default; and Article 32, which lists pseudonymization and encryption among security measures. The EDPB published Guidelines 01/2025 on pseudonymization for public consultation in January 2025 and has since adopted a final version. Its Opinion 28/2024 of December 2024 states that AI models trained on personal data cannot, in all cases, be considered anonymous, and that whether a model is anonymous is assessed case by case. In Case C-413/23 P (EDPS v SRB) of September 4, 2025, the Court of Justice held that pseudonymized data is not personal data in every case and for every recipient: a recipient with no reasonable means of re-identifying the person may hold anonymous data, while the controller that keeps the additional information still holds personal data. The EDPB's draft Guidelines 02/2026 on anonymization, adopted on July 7, 2026 and open for consultation until October 30, 2026, build on that judgment and are meant to update Opinion 05/2014 of the Article 29 Working Party, which remains the main reference on anonymization techniques until the new guidelines are final.

From our projects

In Generator ETR, the app we are building for PSONI, a Polish association for people with intellectual disabilities, users paste or upload documents to have them simplified. Sensitive data is encrypted with Fernet, and content has a default retention of 30 days. In September 2026 our QA confirmed on staging that the scheduled job removes it: a text from August 24 was still available on September 23 and cleared the next day.

An anonymized copy of a text, made with Microsoft Presidio, is created only when the user has not objected to the use of their texts for AI training. The anonymized texts can be exported as one JSON file to build a training corpus. Testing showed that anonymization works but not well enough, so improving it is planned for the next stage of the project; one option under consideration is training HerBERT, a Polish language model, to detect personal data.

Sources

  1. Regulation (EU) 2016/679 (GDPR), Article 4(5) and Recital 26 - EUR-Lex
  2. Opinion 05/2014 on Anonymisation Techniques (WP216) - Article 29 Working Party
  3. Guidelines 01/2025 on Pseudonymisation - European Data Protection Board
  4. Guidelines 02/2026 on Anonymisation, version 1.0 (public consultation) - European Data Protection Board
  5. Judgment of the Court of Justice of September 4, 2025, Case C-413/23 P, EDPS v SRB - Court of Justice of the European Union
  6. Opinion 28/2024 on certain data protection aspects related to the processing of personal data in the context of AI models - European Data Protection Board
  7. Microsoft Presidio documentation - Microsoft

FAQ

Łukasz Gajownik
Łukasz Gajownik
Head of AI & Frontend
Talk to an engineer
  • Usually not. Addresses, ID numbers, dates and context such as a job title or a rare illness can still identify the person, so removing names typically gives pseudonymized data, which is still covered by the GDPR.

  • Truly anonymous data is not (Recital 26). Pseudonymized data is, at least for whoever holds or can reasonably obtain the additional information, and so is any dataset where the person can still be identified with means reasonably likely to be used.

  • Not necessarily for processing, because the text does not go to a model vendor. You still need a retention period, controlled logs and anonymization before texts are reused for training or evaluation.

  • Only with a legal basis for that purpose (for health or other special-category data, usually explicit consent), honoring objections, and after anonymization with a human check of a sample. The EDPB's Opinion 28/2024 states that models trained on personal data cannot, in all cases, be considered anonymous.

Building a system that depends on Data anonymization?

See how we build software for this domain, with case studies and the stack we use.

Requirements checklist

For each term we send the definition and what it requires from your software. Free, no sales call needed.

Your checklist is empty. Use the plus next to a term to add it.

    Order Group sp. z o.o. (Warsaw) uses your e-mail to send the checklist (Art. 6(1)(b) GDPR) and keeps a record of the request (Art. 6(1)(f) GDPR). Marketing consent is optional and can be withdrawn at any time. Read the Privacy Policy