EU AI regulation / data protection

EU Regulators Rule Out Consent for AI Web Scraping, Betting Compliance on a Balancing Test

The EDPB's draft Guidelines 03/2026 force AI developers onto Article 6(1)(f)'s legitimate-interest test and treat robots.txt as legally relevant.

EDPB's New Web Scraping Rulebook for AI People of Internet Research · EU Oct 30, 2026 Public consultation deadline Guidelines 03/2026 open for feedba… 22 pages Length of new guidance First EDPB framework specifically … 3-part test Legitimate interest test prongs Legitimate interest, necessity, an… Sept 4, 2025 CJEU precedent cited Guidelines lean on the CJEU's EDPS… peopleofinternet.com
EDPB's New Web Scraping Rulebook for A… People of Internet Research · EU Oct 30, 2026 Public consultation dea… 22 pages Length of new guidance 3-part test Legitimate interest test pr… Sept 4, 2025 CJEU precedent cited peopleofinternet.com

Key Takeaways

A First Rulebook for Scraping the Open Web

On 8 July 2026, the European Data Protection Board announced it had adopted Guidelines 03/2026 on web scraping in the context of generative AI, alongside final guidelines on anonymisation and blockchain, during its latest plenary session (EDPB news release). The 22-page document, version 1.0, is the first comprehensive attempt by an EU regulator to spell out how the GDPR applies when a company crawls public registers, news sites, forums and personal blogs to build training data for large language models (ppc.land). It is open for public consultation until 30 October 2026 (EDPB consultation page).

The headline move: consent, the legal basis most companies reach for by default, is effectively taken off the table. The guidelines state that consent "will most probably not serve as a workable legal basis for scraping" at internet scale — there is no realistic mechanism for a crawler to obtain informed, freely-given consent from millions of individuals whose data sits on someone else's website. Instead, AI developers are pushed toward Article 6(1)(f)'s legitimate-interest test: a three-part showing that a legitimate interest exists, that scraping is necessary to serve it, and that a balancing exercise doesn't tip in favour of data subjects' rights.

Steelmanning the Board

The EDPB has a real problem to solve, and it deserves a fair hearing before any pushback. Consent-by-default has always been a fiction in mass data collection — nobody meaningfully consents when a training crawler indexes their five-year-old blog comment. Worse, once personal data is baked into model weights, there is no reliable way to delete it; machine unlearning remains immature and expensive. That asymmetry — easy to scrape, nearly impossible to unwind — is a legitimate reason to demand a genuine ex-ante balancing test rather than a rubber-stamped checkbox. The Board's emphasis on data minimisation, source reliability and timestamping before training data ever touches a model is a sensible response to a genuinely hard technical constraint, not regulatory overreach for its own sake.

The guidelines also lean on real jurisprudence rather than inventing standards from scratch. They cite the Court of Justice's 4 September 2025 ruling in C-413/23 P, EDPS v SRB, which clarified that whether data is "anonymous" depends on the recipient's actual ability to re-identify someone — a contextual test the EDPB is now importing into the scraping context.

"Data is anonymous if it does not relate to an identified or identifiable natural person. Whether this is the case may vary from one entity to another." — Anu Talus, EDPB Chair

Where the Balance Tips Too Far

The trouble is in the execution. The guidelines elevate robots.txt, ai.txt files and CAPTCHAs to genuine legal relevance: while "the mere absence of a robots.txt file does not amount to consent," its presence now factors into whether a data subject could "reasonably expect" their data to be scraped — and thus into the necessity and balancing legs of the Article 6(1)(f) test. That converts a voluntary, technically inconsistent web standard with no enforcement mechanism into a quasi-legal signal, without Parliament ever legislating it as one.

The compliance burden this creates is not evenly distributed. As ReedSmith's analysis notes, the guidelines effectively require controllers to maintain searchable records of scraped domains, collection dates, and pre-collection filtering and pseudonymisation pipelines across an entire training corpus (ReedSmith). Large incumbents can absorb that infrastructure cost; university labs, open-source model builders and startups training smaller models cannot as easily. A framework meant to discipline surveillance-scale scraping risks instead entrenching the firms with the deepest compliance budgets — the opposite of the innovation diversity the EU says it wants in the AI market.

There is also a timing problem. The guidelines won't be finalised until after the 30 October 2026 consultation closes, meaning developers must architect data pipelines now against a standard that could still shift. Given that GDPR interpretation tends to harden — not soften — once adopted, few companies will bet on the current draft loosening.

What Should Change Before Finalisation

The EDPB is right that unlimited scraping with no legal basis test is untenable. But a workable rule needs to be as legible to a five-person startup as to a frontier lab — right now, this one is not.

Sources & Citations

  1. EDPB: guidelines on anonymisation and web scraping adopted
  2. EDPB public consultation: Guidelines 03/2026 on web scraping
  3. ppc.land: EDPB blocks AI firms from using consent as an excuse to scrape
  4. ReedSmith: EDPB web scraping guidelines for AI