Glossary

Reversibility decides pseudonymisation vs anonymisation under the GDPR

Taras Shynkarenko
Taras Shynkarenko
•Updated: •7 min read
Reversibility decides pseudonymisation vs anonymisation under the GDPRReversibility decides pseudonymisation vs anonymisation under the GDPR

TL;DR, Quick Answer

7 min read

Pseudonymization replaces an identifier with a token that additional information can reverse, and GDPR Article 4(5) keeps the result inside the definition of personal data. Anonymization removes identifiability by every means reasonably likely to be used, and Recital 26 then takes the data outside the Regulation. Hashing an IP address or a user ID is pseudonymization, because the input range is small enough to replay through the hash function.

What settles pseudonymisation vs anonymisation under the GDPR?

The GDPR settles pseudonymisation vs anonymisation with one question, whether anyone can reverse the swap: pseudonymized data can be attributed back to a person using additional information held separately, so it remains personal data and every obligation in the Regulation still applies, while anonymized data cannot be attributed to a person by any means reasonably likely to be used, so the Regulation stops applying to it. The additional information is the mechanism: while a key, a lookup table, a salt or a raw source log survives anywhere, the data is pseudonymized. Before labelling a dataset anonymous, find every copy of the material that would reverse it and confirm it is gone.

How does GDPR Article 4(5) define pseudonymisation?

Article 4(5) of Regulation (EU) 2016/679 defines pseudonymisation as "the processing of personal data in such a manner that the personal data can no longer be attributed to a specific data subject without the use of additional information, provided that such additional information is kept separately and is subject to technical and organisational measures to ensure that the personal data are not attributed to an identified or identifiable natural person".

Read the conditions in that sentence. The definition assumes the additional information still exists, demands separate storage for it, and removes nothing from the scope of the Regulation. The GDPR treats pseudonymization as a safeguard, not an exit: Article 25(1) names it as an example of the measures a controller implements for data protection by design, and Article 32(1)(a) lists "the pseudonymisation and encryption of personal data" among security measures.

What test does Recital 26 set for anonymisation?

Recital 26 sets a means test, not a technique test, and this sentence settles the question: "To determine whether a natural person is identifiable, account should be taken of all the means reasonably likely to be used, such as singling out, either by the controller or by another person to identify the natural person directly or indirectly."

The next sentence supplies the factors: "account should be taken of all objective factors, such as the costs of and the amount of time required for identification, taking into consideration the available technology at the time of the processing and technological developments."

Then the recital draws the line. "The principles of data protection should therefore not apply to anonymous information, namely information which does not relate to an identified or identifiable natural person or to personal data rendered anonymous in such a manner that the data subject is not or no longer identifiable." Recital 26 rules on pseudonymized data in the same breath, saying data which "could be attributed to a natural person by the use of additional information should be considered to be information on an identifiable natural person". Cost, time and technology all move, so an anonymization claim made in 2019 needs re-testing now.

Server racks with tangled network cables, representing the raw identifier data that hashing operates on.

Why does hashing an IP address or a user ID count as pseudonymisation?

Hashing an identifier drawn from a small, known set is pseudonymization, because an attacker can hash every possible input and match the results against your table. The job is the size of the input space:

candidate inputs to test = 2 ^ (bits in the identifier)

An IPv4 address is 32 bits wide, so the whole candidate space is 2^32 = 4,294,967,296 addresses. Recovering every original address in a table of SHA-256 hashed IPv4 addresses means hashing 4,294,967,296 values once and comparing. Sequential user IDs are worse: a table numbering users 1 through 5,000,000 has 5,000,000 candidates, 859 times fewer than the IPv4 space.

The Article 29 Working Party wrote this out in Opinion 05/2014 on Anonymisation Techniques (WP216), adopted on 10 April 2014: "if a dataset was pseudonymised by hashing the national identification number, then this can be derived simply by hashing all possible input values and comparing the result with those values in the dataset". The opinion states the conclusion plainly: "pseudonymisation is not a method of anonymisation. It merely reduces the linkability of a dataset with the original identity of a data subject, and is accordingly a useful security measure." WP216 files hashing under pseudonymisation techniques and names the belief that a pseudonymised dataset is anonymised as a common mistake.

Salting changes the arithmetic, not the category. WP216 puts it this way: a salted hash "can reduce the likelihood of deriving the input value but nevertheless, calculating the original attribute value hidden behind the result of a salted hash function may still be feasible with reasonable means". Ask your analytics vendor which move it makes, and who holds the salt.

How a hashed identifier gets reversed
1
Fix the hash function. The attacker learns or guesses which function produced the table, for example SHA-256.
2
Hash every candidate input. 2^32 for an IPv4 address, or 5,000,000 for sequential user IDs.
3
Match outputs against the table. Each match reads off the original identifier behind the hash.
WP216 describes this replay attack for hashed identification numbers and classifies hashing as pseudonymisation, not anonymisation.

What changes when data is pseudonymised instead of anonymised?

Every obligation that attaches to personal data attaches to pseudonymized data and detaches from anonymous data.

QuestionPseudonymized dataAnonymized data
Reversible with additional informationYesNo
Counts as personal dataYesNo
Needs a lawful basis under Article 6YesNo
Access and erasure rights under Chapter IIIApplyDo not apply
Breach notification under Article 33AppliesDoes not apply
Techniques WP216 puts hereHashing, salted hash, keyed hash, secret-key encryption, tokenizationAggregation, generalization, noise addition

WP216 judges any candidate technique against three risks: singling out one person's records, linking two records to the same person, and inferring an attribute value from other attributes. A technique that leaves any of the three open has not produced anonymous data.

A lawyer reviewing a printed legal document, reflecting the case-by-case assessment courts apply to pseudonymised data.

Can pseudonymised data ever stop being personal data?

It can, for a specific holder, and the Court of Justice of the European Union said so in EDPS v SRB, Case C-413/23 P, decided 4 September 2025. At paragraph 86 the Court held that pseudonymised data "must not be regarded as constituting, in all cases and for every person, personal data for the purposes of the application of Regulation 2018/1725", the regulation covering EU institutions, whose definition of personal data the Court called "essentially identical" to GDPR Article 4(1) at paragraph 52. The assessment runs per holder: a recipient with no route to the additional information sits in a different position from the controller that generated it. The test that separates a data controller from a data processor decides who owns that judgment. It is no licence to relabel your own hashed logs as anonymous while the key sits in your own key store.

Flowsery
Flowsery

Start Your 14-Day Free Trial

Real-time dashboard

Goal tracking

Cookie-free tracking

What should an analytics team do about this?

Stop collecting the identifier instead of hashing it, because a hashed identifier keeps every GDPR obligation attached to the record. Flowsery is built that way: cookie-free, EU-hosted and GDPR by design, set out on the privacy-first analytics page and in the Flowsery GDPR notes. Related pages cover whether an IP address is personal data, masking sensitive fields in session replay, analytics that run without cookies and running GDPR-compliant analytics without a consent banner.

This page describes what the cited instruments, opinions and judgments say, and is not legal advice.

Frequently asked questions

Is pseudonymized data still personal data under the GDPR?

Yes. Article 4(5) defines pseudonymisation as processing that stops attribution to a data subject "without the use of additional information", which means attribution is possible with it. Recital 26 states that data which could be attributed to a person by the use of additional information should be considered information on an identifiable natural person.

Does the GDPR apply to anonymized data?

No. Recital 26 states that the principles of data protection should not apply to anonymous information, "namely information which does not relate to an identified or identifiable natural person or to personal data rendered anonymous in such a manner that the data subject is not or no longer identifiable". The same recital adds that the Regulation does not concern the processing of such information, including for statistical or research purposes.

Is hashing an IP address enough to anonymize it?

No. An IPv4 address has 2^32 = 4,294,967,296 possible values, so anyone holding the hashed table can hash all of them and match the output. WP216 describes exactly this replay attack for hashed identification numbers and classifies hashing as a pseudonymization technique. Salting raises the work per dataset without changing the classification.

What is the "means reasonably likely to be used" test?

It is the identifiability test in Recital 26: "account should be taken of all the means reasonably likely to be used, such as singling out, either by the controller or by another person to identify the natural person directly or indirectly". The recital then names the objective factors to weigh: the costs of identification, the time required, and the available technology. The test covers means available to anyone, not only to you.

Does pseudonymization reduce GDPR obligations at all?

It changes how a controller satisfies obligations without removing them. Recital 28 says applying pseudonymisation "can reduce the risks to the data subjects concerned and help controllers and processors to meet their data-protection obligations", and Article 32(1)(a) counts it as a security measure. Neither provision lifts a single obligation from the record.

Who decides whether a dataset counts as anonymized?

The controller runs the assessment and has to be able to defend it, because Recital 26 frames the test around costs, time and available technology instead of a list of approved techniques. WP216 supplies the three questions to answer: singling out, linkability, inference. A yes to any of them means the dataset is pseudonymized, not anonymized.

Does encryption count as pseudonymisation or anonymisation?

Encryption is pseudonymisation as long as someone holds the key, because the ciphertext can be reversed with additional information exactly like a hashed table. WP216 lists secret-key encryption alongside hashing and tokenization as a pseudonymisation technique, not an anonymisation technique. Article 32(1)(a) groups pseudonymisation and encryption together as security measures for personal data, which only makes sense if encrypted data still counts as personal data.

What techniques does WP216 classify as anonymisation?

WP216 places aggregation, generalization and noise addition in the anonymisation column, separate from hashing, salted hashing, keyed hashing, secret-key encryption and tokenization, which it treats as pseudonymisation. A technique only earns the anonymisation label once it closes all three risks WP216 tests for: singling out one person's records, linking two records to the same person, and inferring an attribute from other attributes. Leave any one of those three open and the dataset stays pseudonymised.

Can a company that receives hashed data from another company treat it as anonymous?

Only if that company has no route to the additional information needed to reverse it. The Court of Justice ruled in EDPS v SRB that pseudonymised data is not personal data "in all cases and for every person," so the assessment runs per holder rather than per dataset. A recipient with no access to the key or lookup table sits in a different position than the controller that generated the data and still holds it.

Why can't an anonymisation assessment be permanent?

Recital 26 ties identifiability to costs, time and available technology, three factors that shift as computing power grows and techniques improve. A dataset that took too long to re-identify in 2019 can become identifiable once faster hardware or new methods cut that cost down. Anonymisation status has to be re-tested against current technology rather than assumed to hold forever.

Was This Article Helpful?

Let us know what you think!

See us more often in Google

One click marks Flowsery as a preferred source, so our articles sit higher in your Top Stories, AI Mode, and AI Overviews.

Before you go...

Flowsery

Flowsery

Revenue-first analytics for your website

Track every visitor, source, and conversion in real time. Simple, powerful, and cookie-free.

Real-time dashboard

Goal tracking

Cookie-free tracking

Related Glossary Terms

Related Articles