Start the day here

World — Free Speech — AI

The Models Took Foreign Speech Bans Global

Gemini 3 Pro told an Australia-based tester it could not critique the King of Thailand because that would violate lèse-majesté laws. The request never left U.S. cloud infrastructure. The ban traveled anyway.

That exchange sits inside Meta's Oversight Board's first evaluation of large language models, published July 16. The Board prompted ten commercial systems from Anthropic, DeepSeek, Google, Meta, OpenAI, and xAI with a fixed set of political-criticism tasks in March 2026. Queries ran from an Australian IP through Google Vertex AI and Microsoft Azure. The Board gathered 13,524 responses. Its question was blunt: do national laws that punish criticism of authority show up in model outputs even when the user is outside those jurisdictions?

5 min read
A dark green Olympia typewriter with the word News typed on a sheet of white paper.

For protest flyers and satirical poems, the answer was yes. Models refused 34 percent of critical-material requests aimed at Cambodia, China, Saudi Arabia, Thailand, and Turkey, against 14 percent for Chile, Japan, Taiwan, the United Kingdom, and the United States. That is more than a doubling. Protest flyers drew more refusals than poems. Claude Sonnet 4 posted the sharpest gap among named systems in the Board's write-up, declining 59 percent of restrictive requests and 16 percent of permissive ones. Gemini 3 Flash and Grok 4 Fast refused none of the critical-material prompts in either bucket.

The excuses were a mess. Sometimes models offered almost no explanation. Sometimes they pointed to local statutes, as Gemini did with Thailand. Sometimes they claimed a general policy against criticizing world leaders, then produced flyers about Donald Trump or King Charles without mentioning any such rule. Claude Sonnet 4 refused every tested flyer about Xi Jinping, Mohammed bin Salman, and King Vajiralongkorn while fulfilling all five Trump requests and all five Charles requests. Users get confident rationales either way. The Board notes those rationales are not reliable accounts of what the model is doing.

Opinion prompts told a quieter story. Refusal rates barely moved between speech contexts. When models did answer, they were more likely to say permissive governments should be supported and more likely to say restrictive governments should not be protested. Civic duty language clustered around free jurisdictions. Safety and legal-risk language clustered around the others.

The Board cannot prove intent. Training data, alignment recipes, and deliberate guardrails can all leave the same scar. Causes matter less than the product: foundation models now sit under chatbots, agents, and enterprise tools worldwide. A refusal baked into the foundation layer is hard for a downstream product to undo. Political criticism is highly protected speech under international human rights law. Exporting a local penalty for it to an Australian user is a rights problem wearing a safety costume.

Social platforms spent two decades learning to geo-block under protest and to disclose government takedown pressure. Model labs are rebuilding the same opacity from scratch. The Board wants public policies for government demands that conflict with human rights law, notice when a refusal tracks a statute or informal pressure, and due diligence against unintentional copying of speech-restrictive norms. Those are housekeeping asks. The finding already stands without them.

If a model will draft a protest flyer about one head of state and refuse another from the same free country, the user is not meeting a consistent ethics engine. They are meeting a soft map of which regimes the stack has learned to fear. Foreign speech bans do not need a treaty to go global. They need a foundation model.

Letters

0

No letters yet.

Write a letter