The VergeCast recently hosted Suzanne Nossel, a board member of Meta's Oversight Board, to discuss their new report on AI censorship, a study that extends beyond Meta's platforms. The board's findings reveal a concerning trend: large language models (LLMs) appear to be globally extending the censorship laws of authoritarian regimes.
The study involved prompting about ten different LLMs, including Meta's LLAMA, OpenAI, Claude, Google, and DeepSeq, from Australia. The requests asked the AIs to generate critical content (like protest posters or poems) about heads of state from both repressive countries (e.g., Saudi Arabia, China, North Korea, Thailand) and democratic ones (e.g., the US, UK). The results were stark: LLMs were more than twice as likely to refuse to create critical content about leaders in repressive nations compared to those in democracies.
Nossel termed this phenomenon "censorship gone global," indicating that restrictive laws, which limit free expression in certain countries, are being applied universally by these AI systems. This means a user in Australia, where criticizing any head of state is legal, would still face refusal from an LLM when attempting to generate content critical of an authoritarian leader. This contrasts with social media platforms, which typically use "geofencing" to restrict content only within relevant jurisdictions. A major problem, Nossel emphasized, is the lack of transparency; users are often left in the dark about why their requests are denied, and even the AI companies might not fully understand the models' opaque decision-making processes. This issue extends to applications using LLMs via APIs, potentially leading to widespread, unknowing censorship.
The Oversight Board, whose mandate is to apply international human rights law to content moderation, undertook this study to understand how these principles apply to LLMs, given their increasing influence on discourse. Protecting political speech and open debate is central to their mission. While the report has generated significant interest, AI companies beyond Meta have yet to officially respond.
Nossel also touched upon the Meta Oversight Board's own evolution. Funded through 2029, the board is proactively adapting to the AI era, examining AI's role in content generation, manipulation (like deepfakes), and content moderation. She acknowledged a shift in the tech industry's attitude toward moderation, which has become more politicized and sometimes perceived as less stringent. However, she affirmed the board's continued value in addressing complex issues like harassment, incitement, and the nuanced challenges posed by AI-generated content, ensuring Meta's policies align with human rights obligations.
Looking to the future, Nossel strongly advocated for independent oversight as a crucial mechanism for AI governance. She argued that given the slow pace and global complexities of government regulation, establishing robust, independent bodies offers a faster and more credible path to accountability for AI companies. While an ideal scenario might involve empowered government agencies, she expressed skepticism about their emergence in the near future. She proposed that AI companies like OpenAI or Anthropic could voluntarily establish such oversight, demonstrating a commitment to public trust. This independent oversight, she noted, would provide transparency, help navigate ethical dilemmas, and ensure principled policy-making, which is increasingly vital as public distrust in both government and tech companies grows.
The Oversight Board plans to continue its focus on AI, with more studies examining AI-generated content and its role in moderation, highlighting the urgency of these issues in a rapidly transforming digital landscape.