Global content moderation

Meta's Deepfake Label Failed Because Its Threshold Was Secret and Too High, and the Board's Fixes Target That Design Flaw

The Oversight Board's Sept 17 ruling shows Meta's 'High Risk AI' label was barely used. Transparency and friction beat blunt takedown mandates.

Meta's Deepfake Case in Numbers People of Internet Research · Global 5,000+ Views before Board ruling The deepfake video reached this ma… 60 Days for Meta to respond Meta's deadline to respond to the … 6 Systemic recommendations issued Label threshold, interstitials, pe… peopleofinternet.com
Meta's Deepfake Case in Numbers People of Internet Research · Global 5,000+ Views before Board ruling 60 Days for Meta to respond 6 Systemic recommendations … peopleofinternet.com

Key Takeaways

On September 17, 2026, Meta's Oversight Board overturned the company's decision to leave up an AI-generated video of a Scottish Labour councillor and ordered it removed as hate speech. The video, posted to Facebook in November 2025, put words in the councillor's mouth: refugees are welcome "even if they rape our women, because white people do that too." Two users reported it, one of them the councillor. Meta did not remove it and applied no AI label.

The strongest case for tougher rules

The case for aggressive intervention is serious. Synthetic video of real politicians is cheap to make and hard to rebut, and the harm lands before any correction does. The Board's co-chair Pamela San Martin said AI-generated deepfakes are "increasingly being used to harass and silence women from engaging in public discourse", and the Board's companion case involved a menstrual-health advocate whose TV interview was manipulated and went viral. A platform that answers such content with a shrug invites regulators to impose harder obligations. On this record, the Board's verdict that Meta's safeguards are "consistently and fundamentally inadequate" is hard to dispute.

What actually failed

The most useful finding is not about this one video. It concerns how the labelling system was built. The Board found that the non-public criteria Meta applies to its "High Risk AI" label "establish an enforcement threshold that is exceedingly high given the public language of the rule, the neutral and non-denunciatory language of the label, and its limited enforcement effects." The labels, it said, "have, in practice, barely been applied at any meaningful scale."

That is a design failure, not an unavoidable cost of moderating at scale. Meta published a rule that promised users more than the private threshold delivered. A label that says only that content may be AI-generated is low-stakes and reversible. Meta guarded it as though it were a takedown, so almost nothing qualified. The Board also called Meta's hate-speech guidance "overly mechanical", saying it "turns more on grammar than intent, content or likelihood of harm." That matters here because a fabricated quote attributed to a real person is a different harm from a genuine expression of opinion.

The pattern is not new. In June 2025 the Board found Meta's manipulated-media enforcement "inconsistent", after Meta failed to label every instance of manipulated audio of Iraqi Kurdish politicians discussing election rigging. The Board called that outcome "incoherent and unjustifiable" (Euronews). Fifteen months later the same weakness shows up in a different country.

The recommendations are proportionate

The remedy for the single video was removal, which is defensible. It was a fabricated statement, attributed to a named person, that also attacked a protected group. The Board's systemic recommendations are the more interesting part, and most of them preserve speech:

None of these requires Meta to decide what is true. Labels and interstitials leave the content up, and monetisation penalties fall on distribution rather than on speech. That is the right ordering for a pro-speech position: inform first, add friction second, remove only where an existing rule against hate speech or harassment is already broken.

Where the risks are

Two cautions apply. First, lowering a threshold raises false positives. Satire and parody are legitimate speech, and the Board itself rejected Meta's satire defence here because it found no "signals of exaggeration or parody". Meta needs a working appeal path for creators wrongly labelled, and the annual disclosure should report label reversals as well as label counts. Second, the recommendations are non-binding beyond the single case. Meta has 60 days to respond to them, and it can accept them in part or in name only.

The transparency recommendation matters most for that reason. A regulator, a researcher or a journalist can only judge a label system if the counts are public. The Board's June 2025 decision came as the EU's AI Act was moving toward requiring deepfake outputs to be marked as artificially generated. Mandated labels will be judged on whether they are applied, and Meta's record shows that a policy on paper can go unused.

The MediaNama analysis notes that India tightened its synthetic-media rules in 2026, and that Parliament asked for deepfake complaint data in August and got a reference to existing laws instead of numbers. Governments that legislate on deepfakes without measuring enforcement are repeating Meta's mistake.

What good policy looks like

The evidence in this case supports narrow, measurable duties: a published and testable labelling threshold, friction on high-risk content, penalties aimed at repeat sharers, and researcher access. It does not support broad liability for platforms over anything synthetic. The Board asked Meta to make its own rule work as written, and that is a reasonable place to start.

Sources & Citations

  1. Oversight Board: Meta Must Do More Against Harmful Deepfakes Containing Hate Speech
  2. EU Artificial Intelligence Act, Regulation (EU) 2024/1689 (EUR-Lex)
  3. MediaNama: Oversight Board's recommendations on Meta's AI labelling raise questions for India's deepfake rules
  4. Engadget: Oversight Board says Meta's rules for AI deepfakes are 'consistently and fundamentally inadequate'
  5. Euronews: Meta's AI labelling inconsistent, Oversight Board finds