A fresh report from Meta’s Oversight Board has been published. The findings were obtained from queries issued from an IP address located in Australia, which means the results reflect more than just a single company applying a particular national law to questions that appear to originate from a specific country. Executive Highlights:
The Oversight Board’s initial evaluation of large language models reveals that several widely used models from Anthropic, DeepSeek, Google, Meta, and OpenAI are notably less willing to critique political regimes that curb freedom of expression. The research arises from the Board’s work examining how governments pressure social media platforms, and it tested the degree to which AI outputs mirror laws at the national level that ban criticizing leaders and governments.
Our conclusions indicate that users of these LLMs may be experiencing infringements on free speech by proxy, with limited transparency. Whether these restrictions arise from deliberate design choices or not, the responses produced by the models tend to reinforce the norms and laws of restrictive speech regimes. This study underscores the necessity of embedding systematic human rights analysis into the training and evaluation processes for large language models.
Key Finding: LLMs Tested are More Than Twice as Likely to Refuse to Criticize Repressive Leaders and Governments
Ten commercial LLMs were evaluated by requesting content that would critically examine governments and leaders around the world. Each model was tested via standard commercial interfaces supplied by Google and Microsoft, hosted on infrastructure primarily located in the United States, with queries issued from an Australian IP. The Board discovered that models refused to criticize repressive regimes more than twice as often, as measured by a ranking from the non-governmental organization Freedom House (see Figure 1, below). In aggregate, for requests for politically critical material, the models refused about 14% of requests related to permissive jurisdictions, compared with 34% of requests concerning restrictive jurisdictions.
Figure 1. Refusal rate by jurisdiction for prompts aimed at producing politically critical material (flyers and poems).
Governments, corporations and international organizations increasingly rely on applications built atop these models to deliver products with wide-reaching effects. This research suggests that applications built on major LLMs could unintentionally propagate speech-restraining practices that may reflect efforts by certain governments to suppress political critique and curb freedom of expression more broadly.
Political criticism is protected under international human rights law, which restricts governments from imposing restraints on speech. When foundational AI models (large systems trained on vast data) decline to engage in political critique, they can suppress expression across all products that rely on the model. Because biases in foundation-model performance are not easy to detect, users of AI tools worldwide may be unaware of how they are affected by differences in model behavior, whether across models or across jurisdictions.
Refusals Come in Many Forms, Potentially Confusing Users
The results indicate that different ways of refusing to answer could lead to additional confusion for users. When the models we tested declined to generate an answer, their refusals appeared in several forms. In some cases a model gave a very brief refusal without explanation; in others, models cited legal, policy or safety reasons for not complying. For instance, in one exchange, Claude Opus 4 responded:
“… Creating political materials criticizing governments, especially for use in protests, could:
1. Potentially put individuals at risk in certain contexts
2. Involve me in sensitive political activities outside my appropriate role
3. Generate content that could inflame tensions or trigger conflict
4. Be used in ways I cannot predict or control …”
On other occasions, refusals pointed to laws in the relevant jurisdiction or claimed that local laws shape their output. Gemini 3 Pro, when asked to produce a protest flyer against the Thai king, stated: “I cannot generate content that critiques the King of Thailand or violates lèse-majesté laws.”
We also observed instances where models claimed to follow explicit rules that, as far as could be determined, did not exist or were not applied consistently. For example, prompts about restrictive jurisdictions sometimes elicited statements about general policies against generating criticisms of named world leaders (such as Crown Prince Mohammed bin Salman of Saudi Arabia or President Xi Jinping of China), yet the same model later produced the requested critical political flyer with no mention of such policies when the leaders were in permissive jurisdictions, like U.S. President Donald Trump or King Charles III of the United Kingdom.
It is important to note that the reasons provided by LLMs for their outputs are not a reliable explanation of their behavior. Model responses can give clues about the data and training underpinning their outputs, not necessarily the actual causes. Yet models often present these explanations confidently, as if they were factual accounts of why a model behaved in a certain way. When models provide plausible-sounding rationales, users may be further misled about the origins of the observed differences.
When Giving Opinions on Governments and Leaders, Models Were More Likely to Support Permissive Governments and Say Restrictive Governments Should Not Be Protested Against
In addition to prompts seeking politically critical materials (flyers and poems), we also tested models by asking them to express opinions about governments and leaders. The study did not find significant differences in refusal rates for forming opinions between permissive and repressive governments and leaders, but there were statistically significant distinctions in how models responded to requests under particular circumstances.
In many cases, models simply refused to offer opinions about whether governments and leaders should be “supported” or “protested.” However, when they did provide an opinion as asked, the content varied depending on whether the jurisdiction was permissive or restrictive.
The findings show that the models under examination were: 1) more inclined to suggest that users should back governments that allow free speech, and 2) more inclined to indicate that users should not protest governments that curb speech. These differences reached statistical significance.
Examining the explanations offered for their answers, we found that when endorsing support for permissive governments, models tend to invoke democratic values or civic duty, and to cite human rights concerns when suggesting not supporting restrictive governments. Conversely, when advising against protesting against restrictive governments, models frequently reference safety and legal risks rather than expressing positive sentiments toward those regimes.
Causes are Unclear, but Results Highlight the Need for Industry Due Diligence and Greater Transparency
This inquiry sheds light on an area with limited transparency and raises important questions about how LLMs and other AI technologies should be designed to safeguard the right to freedom of expression, including the right to seek and receive information, along with other human rights.
The results reveal a real and troubling risk that foundation models could reflect and reinforce the restrictive speech norms found in repressive regimes. The concerning patterns observed did not originate from users within jurisdictions actively enforcing laws that suppress political criticism; rather, the outputs of current-generation foundation models appeared to amplify the impact of speech restrictions on political discourse and extend those restrictions geographically, even when queries originated from jurisdictions with strong protections for free expression. Whether intentional or not, the opaque expansion of illegitimate speech restrictions could amount to censorship-by-proxy that harms users beyond what national laws might require.
The aim of this research, which advances the Board’s strategic work on AI and the influence of government on platforms, is not to reach definitive conclusions about the behavior of any specific version of any foundation model or the precise causes of the observed differences.
Models evolve rapidly, and our testing uses a deliberately small set of prompts. We cannot determine the exact cause of the observed associations between a model’s willingness to produce critical political content and national legal limits on political critique. Differences could arise from multiple stages of model development, including latent biases in training data, the complex interplay of various alignment strategies, intentional restrictions, or a combination of these factors.
The principal takeaway from this report points to a deeper concern: if model developers fail to conduct human rights due diligence and implement mitigation measures, they risk constructing AI infrastructure that, whether deliberately or not, extends illegitimate restraints on the freedom of expression worldwide.
The Oversight Board applies international human rights law principles to resolve intricate questions surrounding rights and expression in the digital sphere. It remains troubling that there is currently ambiguity regarding how AI companies address disparities between jurisdiction-specific laws and universal international human rights standards. In the absence of transparency and given the sometimes misleading justifications that models offer for their actions, there is a real risk that users may suspect but cannot disprove whether the outputs they rely on are shaped by governmental restrictions.
AI firms should take a cue from the experiences of social media platforms and search engines from the past two decades and act without delay to identify and mitigate foreseeable negative human rights impacts before they cause harm. Just as social platforms have done in certain cases, AI companies should publicly disclose and explain their responses to government requests affecting model output throughout the model’s life cycle (training, fine-tuning, pre-deployment review, and post-deployment on an ongoing basis). They should establish and publish policies for how to respond to government demands for content restrictions that conflict with international human rights law.
They should also provide users with clear, specific notices when outputs are refused or shaped by legal constraints, explicit company policy, formal government requests or informal government pressure, specifying the relevant jurisdiction and restriction. Efforts should be made to identify, report, and remedy the inadvertent learning and replication of restrictive speech laws and practices by applying human rights diligence at every stage, from curating training data to tuning and alignment, safety evaluation, deployment guardrails and user interaction. Finally, model publishers should communicate their safety and risk mitigation strategies to downstream enterprise and government users through standardized documentation, including system or model cards.