The EDPB's two July 2026 draft guidelines, on anonymization and on web scraping for generative AI, replace the 2014 re-identification test with three new pass criteria. Here is what the texts say, clause against clause, and how a product-analytics export fares against them.
++++
field notesNº08
fig — When does 'anonymised' data leave GDPR scope?
The EDPB's two July 2026 draft guidelines, on anonymization and on web scraping for generative AI, replace the 2014 re-identification test with three new pass criteria. Here is what the texts say, clause against clause, and how a product-analytics export fares against them.
"We anonymised it" is the most load-bearing sentence in a security questionnaire, and the EU-level test behind it was last specified in April 2014. On 07 July 2026 the EDPB adopted its first restatement in twelve years: draft Guidelines 02/2026 on Anonymisation and draft Guidelines 03/2026 on web scraping in the context of generative AI. Both are Version 1.0, adopted 07 July 2026 for public consultation; per the EDPB's announcement, "the guidelines will be subject to public consultation until 30 October 2026." Every quote from these two documents below is draft text; it can change. I read both against the settled anchors: GDPR Recital 26, WP29 Opinion 05/2014 (WP216), and the EDPB's final Opinion 28/2024 on AI models. The headline finding: draft 02/2026 replaces WP216's three-part test with three renamed and re-scoped criteria, and failing one of them is no longer the end of the analysis.
"'personal data' means any information relating to an identified or identifiable natural person ('data subject'); an identifiable natural person is one who can be identified, directly or indirectly, in particular by reference to an identifier such as a name, an identification number, location data, an online identifier or to one or more factors specific to the [...] identity of that natural person;"
Whether someone is identifiable is decided by Recital 26, the crux clause of everything below, in full:
"The principles of data protection should apply to any information concerning an identified or identifiable natural person. Personal data which have undergone pseudonymisation, which could be attributed to a natural person by the use of additional information should be considered to be information on an identifiable natural person. To determine whether a natural person is identifiable, account should be taken of all the means reasonably likely to be used, such as singling out, either by the controller or by another person to identify the natural person directly or indirectly. To ascertain whether means are reasonably likely to be used to identify the natural person, account should be taken of all objective factors, such as the costs of and the amount of time required for identification, taking into consideration the available technology at the time of the processing and technological developments. The principles of data protection should therefore not apply to anonymous information, namely information which does not relate to an identified or identifiable natural person or to personal data rendered anonymous in such a manner that the data subject is not or no longer identifiable. This Regulation does not therefore concern the processing of such anonymous information, including for statistical or research purposes."
Three details matter here. The standard is objective and priced: "the costs of and the amount of time required for identification", against available technology and its development. The means counted are anyone's: "either by the controller or by another person". And a wording point: the phrase "singled out" appears nowhere in the GDPR. Recital 26 says "such as singling out", and that is the only place the concept appears; Article 4(1) contains neither phrase. "Reasonably likely" itself appears nowhere outside this recital — twice here, and never again in the regulation.
2014's three attacks: singling out, linkability, inference
"– Singling out, which corresponds to the possibility to isolate some or all records which identify an individual in the dataset;– Linkability, which is the ability to link, at least, two records concerning the same data subject or a group of data subjects (either in the same database or in two different databases). [...]– Inference, which is the possibility to deduce, with significant probability, the value of an attribute from the values of a set of other attributes."
Run against your own tables: can any record be isolated down to one person; can records about the same person be joined, within one database or across two; can an attribute be deduced from the others. WP216 ties the trio straight back to the legal standard on p. 12: "a solution against these three risks would be robust against re-identification performed by the most likely and reasonable means the data controller and any third party may employ."
Two operational positions from 2014 still bite. From p. 9: "when a data controller does not delete the original (identifiable) data at event-level, and the data controller hands over part of this dataset (for example after removal or masking of identifiable data), the resulting dataset is still personal data." And from p. 24: "Do not rely on the 'release and forget' approach" — re-evaluate the residual risk regularly, adjust the controls, monitor.
What draft 02/2026 keeps, and what it replaces
Draft Guidelines 02/2026 acknowledges its predecessor in para 2 (this and every 02/2026 quote in this post is draft text): Opinion 05/2014 "provided three criteria for assessing the anonymity of data", but since 2014 "there have been major changes in the legal, privacy, engineering and technological landscapes." Footnote 3 grandfathers existing work: a controller that assessed a dataset under Opinion 05/2014 before publication "is not expected to conduct a new assessment."
Two logged searches: "WP216" appears nowhere in the draft; neither does "linkability". The draft speaks of "Opinion 05/2014" and renames the whole apparatus. From the executive summary (p. 2):
"The framework itself presents three criteria which can be used to test if data is anonymous: No Record Isolation, No Linkage and No Inference."
Side by side, clause against clause:
WP216 (2014, final), pp. 11–12
Draft 02/2026, §3.4 (in consultation)
What moved
"Singling out, which corresponds to the possibility to isolate some or all records which identify an individual in the dataset"
Para 55: "The No Record Isolation criterion is met if the data does not contain a unique combination of attribute values that relate to a single individual."
The attack becomes a structural uniqueness check; singling out itself moves to a follow-up stage (paras 98–99).
"Linkability, which is the ability to link, at least, two records concerning the same data subject or a group of data subjects (either in the same database or in two different databases)."
Para 60: "The No Linkage criterion is met if the data does not contain an individual's record which could be linked to another record which (a) also relates (with certainty or high likelihood) to that same individual, and (b) comes from a different dataset."
Narrowed twice: individual-level only, cross-dataset only, with an explicit likelihood threshold. Group and aggregate linkage migrate to the third criterion (para 80: such linkage "can also be seen as a form of inference").
"Inference, which is the possibility to deduce, with significant probability, the value of an attribute from the values of a set of other attributes."
Para 67: "The No Inference criterion is met if no specific and meaningful inference can be drawn from the given data."
A materiality filter is added: the inference must relate to "a single identified or identifiable individual", be "liable to have an effect on the data subject's rights and interests", and not be obtainable "from general knowledge or from data about the population at large".
The polarity of the whole test changed with it. Para 52 (p. 18):
"violating a criterion does not necessarily mean that the information must necessarily be considered as personal data; rather, if one or more of the criteria is violated, it is then necessary to continue the analysis, as set out in sub-section 3.5.3, to assess the impact of that violation and whether the data could still be considered anonymous."
The 2014 version of that rule, still quoted by the final Opinion 28/2024 from WP216 p. 24: "whenever a proposal does not meet one of the criteria, a thorough evaluation of the identification risks should be performed". The draft turns that footnote into architecture. Para 98 sends a failed No Record Isolation result into a follow-up singling-out analysis, and para 99 states an outcome 2014 never wrote down: "If this singling out cannot be done – and the other two criteria have been successfully met – then the given data can still be considered anonymous, despite failing the No Record Isolation criterion." Unique records can still be anonymous if there is nothing to match them against.
What did not change: it is still a three-part test, and the draft still grounds everything in Recital 26's "reasonably likely" standard (46 occurrences in the text layer). Recital 26's factor list survives inside para 29's expanded version, with one addition: "technological developments" becomes "reasonably foreseeable technological developments". The draft also sharpens identifiability itself (exec summary, p. 2): a person counts as identifiable only if distinguishing them happens "in a way that makes it possible to treat them differently", an added limb relative to Recital 26. And para 31 rules out one mitigating argument: "The EDPB cautions against using 'lack of motivation' as a factor in this analysis."
One fork to watch: the final Opinion 28/2024 (para 40) still anchors model anonymity on the trio under its 2014 names — "single out, link and infer". If draft 02/2026 is finalized as-is, the EDPB's final Opinion and its new guidelines will name the operative test differently. That is a consultation comment waiting to be written.
Scraped data and the model-anonymity question
Settled law first. Opinion 28/2024 (final, adopted 17 December 2024) holds, at para 34, that AI models trained on personal data "cannot, in all cases, be considered anonymous. Instead, the determination of whether an AI model is anonymous should be assessed, based on specific criteria, on a case-by-case basis." The mechanism is para 31: training data "may still remain 'absorbed' in the parameters of the model". Per para 38, a supervisory authority should see "sufficient evidence that, with reasonable means: (i) personal data, related to the training data, cannot be extracted out of the model; and (ii) any output produced when querying the model does not relate to the data subjects whose personal data was used to train the model."
Now the scraping draft. The scope clause of draft Guidelines 03/2026 (para 12, draft text) restates applicability: "The GDPR will apply to web scraping when it includes processing operations, such as extraction, cleaning, structuring and storing of personal data", including "data that indirectly identifies an individual (i.e. information that can be linked to an identifiable individual considering all the means reasonably likely to be used)." One note: the "2020603" in the PDF filename is the EDPB's own typo; the link is cited as published.
What the draft does not do, established by logged searches rather than assumption: it sets no conditions under which scraped data or a resulting model leaves GDPR scope. It never mentions Recital 26, never cites Opinion 05/2014, and never cites its sibling draft 02/2026. Anonymization appears only as a mitigation measure — para 38 (draft text): "where feasible, anonymise or alternatively pseudonymise personal data"; para 66(g): "Deleting or anonymising the unnecessary personal data as soon as possible".
The EDPB's news page says the scraping guidelines build on the AI-models Opinion. Checked against the PDF, that claim is accurate and precisely scoped: all eight citations of Opinion 28/2024 sit in the purpose-limitation and legitimate-interest sections. So the three documents divide the labor. Draft 02/2026 says when data is anonymous. Final Opinion 28/2024 says when a model is anonymous. Draft 03/2026 says how to scrape lawfully assuming the GDPR applies — including, per para 14 (draft text), that "compliance with regard to Article 5(2) GDPR must be met by the controller" even where the draft concedes accountability is "not obvious". The seam I would raise in a consultation response: 03/2026 recommends anonymizing scraped personal data without naming the test the anonymization must pass, and without citing 02/2026.
Enforcement context: Italy's Garante already addressed large-scale scraping by generative-AI developers in its provision of 20 May 2024 (registro n. 329, Gazzetta Ufficiale n. 132 of 7 June 2024; Italian only), including defense measures site operators can deploy.
Running the three criteria against a product-analytics export
The position, stated exactly as far as the texts carry it: an export that is (i) record-level, (ii) keyed by a persistent identifier, and (iii) high-dimensional sits on the failing side of the draft's own criteria. Whether that describes "most" analytics data, no text says, and I won't either.
The draft's rule of thumb (para 86): "As a rule of thumb, (re-) identification techniques are more likely to be effective against record-level data with high dimensionality and high resolution. By contrast, many (re-)identification techniques are less likely to be effective against data which is aggregated through mathematical functions into statistical indicators (e.g. averages)". Para 85 defines dimensionality as "the number of columns in a table for each individual record"; "a full date of birth has a higher resolution than the year of birth alone". A product-analytics event table scores high on both by design.
On persistent identifiers, para 62 (draft text): "Linkage may often be possible by reference to a common identifier (e.g., an internal customer number) which is used in both datasets, or by matching records based on certain combinations of attributes." And the draft's Example 6 is the web-analytics case itself: several websites share page-visit data with a third-party ad provider, keyed by "a pseudonym which is based on fingerprints from the users' devices". The draft's conclusion: "both the raw fingerprint data (as a combination of attributes) and the pseudonym which is derived from that data (as an identifier) allow for the identification of the individual. This is true for both the websites and the third party." Cost is no refuge either; per para 92, most documented re-identification techniques "require limited resources and can be performed within a reasonable timeframe, even on commodity hardware", especially against record-level data.
Hashing or rotating the key does not change the classification by itself. Draft 02/2026 para 50 lists among the immediately-not-anonymous cases a dataset with "pseudonyms that can easily be reverse engineered to discover an identifier". The EDPB's pseudonymization guidelines (Guidelines 01/2025, still a consultation version adopted 16 January 2025 with no final text as of this writing; draft 03/2026 itself cites it with that status spelled out) say in para 22 that even after erasing all additional information, "the pseudonymised data becomes anonymous only if the conditions for anonymity are met."
The escape hatches are real, though. Para 99 (quoted above) lets unique records stay anonymous when nothing can be matched against them, and Example 24 shows it: a secret-ballot cinema survey where "there is no way to link the answers to their identity ... nor to treat any respondent differently" is anonymous despite failing No Record Isolation. Those hatches do little for a keyed analytics export: the persistent identifier exists precisely to link the same person across sessions and tables. It supplies the matching capability Example 24's respondents lack.
The check itself, ordered the way the draft orders it. Every line traces to a quoted clause:
shell
Re-identification check — <export name>, <date>Perspective: every entity that can access the data(draft 02/2026 para 30: controller, receiving entities,"Rogue employees", malicious actors; "lack of motivation"is not a factor, para 31).[ ] 1. Record isolation (para 55) Does any combination of attribute values relate to a single individual? A persistent user ID or device fingerprint does, on Example 6's reasoning; more columns raise uniqueness (para 56).[ ] 2. Follow-up singling out (paras 98-99) If records are isolated: can they be matched to a person's attributes in any data available to the entities above? If not, and the other criteria hold, the data can still be anonymous (para 99).[ ] 3. Linkage (paras 60, 62)
The first item has an obvious shape in SQL (illustrative; substitute your own quasi-identifier columns and table):
-- Unique combinations violate No Record Isolation (para 55)
SELECT plan, country, signup_week, COUNT(*) AS matches
FROM analytics_export
GROUP BY plan, country, signup_week
HAVING COUNT(*) = 1;
The evidence you'd have to keep, and what to write in the questionnaire
Where each artifact is grounded matters here, because two of the three governing documents can still change:
Anonymization process documentation, including test runs — the only artifact draft 02/2026 itself names. Para 41 (draft text): "controllers should ensure adequate documentation of the anonymisation processing. This documentation of the anonymisation process (including for the testing of the supposedly anonymous datasets) makes it possible to demonstrate both the GDPR-compliance of the anonymisation process, and the effectiveness of the anonymisation itself." Logged searches: zero occurrences of "DPIA", "Article 30" or "accountability" in the draft.
DPIA and records of processing — grounded in the final Opinion 28/2024, not the drafts. Para 56: "Articles 5, 24, 25, and 30 GDPR, and, in cases of likely high risk to the rights and freedoms of data subjects, Article 35 GDPR, require controllers to adequately document their processing operations", and that holds "even if the objective of the processing is anonymisation." Para 58(a) includes "any information relating to DPIAs, including any assessments and decisions that determined that a DPIA was not necessary". Draft 03/2026 para 34 (draft text) adds "undertaking a data protection impact assessment and making it publicly available" as an appropriate measure.
Test reports — final Opinion 28/2024: para 55 names structured testing against attribute and membership inference, exfiltration, regurgitation of training data, model inversion and reconstruction attacks, and para 58(e) asks for "reports on how the model has been tested (by whom, when, how and to which extent)" plus the results.
The consequence of not keeping any of this — final Opinion 28/2024 para 57: an authority that cannot confirm anonymization from the documentation "would be in a position to consider that the controller has failed to meet its accountability obligations under Article 5(2) GDPR."
Scraping-specific artifacts — draft 03/2026: documented instructions to a scraper acting as processor, "in particular on how the data set should be constituted with regard to data sources and categories" (para 18); precise collection criteria and "a data mapping and inventory process" (para 37); and scrape timestamps — "record when the data was scraped, to allow to show how up to date the data is" (p. 13).
A legal basis for the anonymization step itself — draft 02/2026 para 38: "anonymisation must have a legal basis under Article 6 GDPR", with an Article 9(2) exemption where special categories are involved.
One more 02/2026 line (draft, para 40) belongs next to every questionnaire answer: "controllers should not use descriptions like 'anonymous', 'de-identified' or 'de-personalised' if individuals are actually still identifiable."
Put together, an answer the texts support names a test, a perspective, a date, and a document location:
shell
Q: Is the data you provide/retain anonymised?A: We classify <dataset> as <anonymised | pseudonymised, i.e. still personal data>. Assessed against the re-identification criteria in WP29 Opinion 05/2014 pp. 11-12 and EDPB draft Guidelines 02/2026 section 3.4 (No Record Isolation, No Linkage, No Inference — draft, in public consultation until 30 October 2026). Perspective: <entities that can access the data>. Result: <per criterion, incl. the follow-up singling-out analysis for any failed criterion>. Last assessed: <date>; re-run on new data sources, new techniques, security incidents (draft 02/2026 paras 42, 84; WP216 p. 24: no "release and forget"). Evidence: anonymisation process record incl. test reports at <location> (draft 02/2026 para 41; Opinion 28/2024 paras 56-58); DPIA entry or documented decision that none was needed (Opinion 28/2024 para 58(a)).
The test behind "we anonymised it" exists in writing again, in draft. Run the three criteria against your own exports before a questionnaire asks. If the draft wording would break your answer, the consultation on both guidelines is open until 30 October 2026.
Keeping that evidence trail current (the anonymization process record, the test runs, the DPIA entries) is the sort of thing devguard exists to make routine.
// authored by
VS
Vadim Sikora
Co-founder · devguard
Vadim leads engineering, the architecture, automation, and integrations that let compliance live. He's set on making a serious, audit-grade system genuinely easy to run.