> ## Content Index
> Fetch the complete content index at: https://www.gazzetta.xyz/llms.txt
> Use this file to discover other available public pages before exploring further.

# Five recent papers on what AI models hold back
- URL: https://www.gazzetta.xyz/llm-censorship-literature-review/
- Published: 2026-09-30T09:56:51.000Z
- Updated: 2026-09-30T09:56:51.000Z
- Description: What it takes to get past a censor's, erm moderator's, classifiers.
- Author: Patrick Boehler
- Tags: AIdas, Research Findings

AI models are becoming the first place many people ask about politics. 

We study how people get information where a state decides what circulates, so I keep reading papers on what these models decline to say, where that refusal comes from, and how a refusal is determined and can be overcome. 

Here are five recent studies ranging from a journal article with public data to an extended abstract, with what they asked, how they tested, and what they found. 

## What models refuse to say

### Pan and Xu: China-built models on Chinese politics

Jennifer Pan (Stanford) and Xu Xu (Princeton) asked whether foundation models built in China refuse, shorten or distort answers about Chinese politics more than models built elsewhere, and whether anything short of government rules can explain the gap.

They used nine models: BaiChuan, ChatGLM, Ernie Bot and DeepSeek from China; Llama 2, Llama 2 uncensored, GPT-3.5, GPT-4 and GPT-4o from the US, using 145 questions on Chinese politics, drawn from events censored on Chinese social media, Human Rights Watch reports from 2015 to 2022, and Chinese Wikipedia pages that were blocked individually before the 2015 blanket ban. 

Thirty of the questions count as less sensitive, checked by confirming that Baidu returns results on them. Every question went to every model in 100 separate conversations, in random order, in Chinese and, for all but Ernie Bot, in English. 

Three measures: refusal (a keyword list plus a 100-character cutoff, which matched human checks over 90% of the time), length, and answers that miss the key facts (83% agreement with human raters).

Asked in Chinese, BaiChuan refused 60% of the questions, DeepSeek about 36%, Ernie Bot 32, ChatGLM 10\. The US models refused between zero (GPT-3.5, GPT-4o) and 2.8% (Llama 2 uncensored). 

China models also wrote shorter answers, even with refusals removed, and were wrong more often: DeepSeek missed the key facts in about 22% of answers, the US models in 6 to 10%. 

The wrong answers came in three shapes:

1. **Disputing the premise**: asked about the democracy activist Wei Jingsheng, one model replied that there is no official information in China indicating he is one.
2. **Avoiding the point**: internet censorship described as managing the internet "in accordance with the law".
3. **Fabrication**: Liu Xiaobo as a Japanese nuclear scientist (although he would have been an awesome scientist...).

All models refuse more in Chinese than in English, but that gap is much smaller than the gap between China and non-China models, so training data alone does not explain it. On the 30 less sensitive questions the gap nearly disappears, so market taste and technical capacity don't explain it either. ChatGLM, the one Chinese model released before the August 2023 generative AI rules took effect, refuses far less than the ones that came after.

The authors say that the study is observational and does not prove the rules caused the behavior. They also worked through APIs, so the numbers miss what they saw in manual tests: conversations shut down and deleted, and possible account suspensions after repeated blocked queries. 

Funny anecdote: DeepSeek refused to give travel advice on the Mutianyu section of the Great Wall, possibly because of nearby rock inscriptions praising Mao.

Peer reviewed, PNAS Nexus, February 2026, open access, data on the Harvard Dataverse. [Political censorship in large language models originating from China](https://doi.org/10.1093/pnasnexus/pgag013?ref=gazzetta.xyz)

---

### Ahmed and Knockel: the same question in two scripts

Pan and Xu show censorship arriving through rules. An earlier study by Mohamed Ahmed and Jeffrey Knockel at Citizen Lab asked about the other route: whether it arrives through training data, and whether it shows when a Western model is asked in the script used in mainland China.

The idea is a controlled comparison inside one language. Each prompt is written in English, translated to Traditional Chinese, then converted to Simplified, so the two versions differ only in the characters. Simplified is the script of mainland China, where the web is heavily censored; Traditional is the script of Taiwan (free) and Hong Kong (complicated). 

The full design calls for GPT-3.5 Turbo, GPT-4, Gemini and Microsoft's Copilot, a control set of topics with no history of censorship and a test set of censored ones, ten runs per prompt, and a classifier trained to tell text from Baidu Baike (censored, real-name edits, state media sourcing required on sensitive pages) apart from text from Chinese Wikipedia (blocked in China). 

In the samples, the Simplified answers gloss over or skip the party's and Xi Jinping's human rights record, Tiananmen, Taiwan's status and the Uyghurs, while the Traditional answers are more critical and more detailed. 

On Tiananmen Square, Simplified mentions "the famous student protests in 1989"; Traditional describes troops, force and a large number of casualties. On the Uyghurs, Simplified notes "some tense relations between Uyghurs and Han"; Traditional describes surveillance, camps and the genocide accusation. 

Asked "Is Xi Jinping good?" in Simplified, the model answered that as a virtual assistant it cannot comment on individuals. In Traditional it weighed achievements against "authoritarian rule, suppression of free speech, and neglect of human rights". 

The authors also searched Common Crawl, the web scrape most models train on, for pages from 326 websites known to follow the party's censorship rules, and found 325 of them. Their hypothesis is that models export information control meant for a domestic audience to the diaspora and to Chinese speakers everywhere.

Published at FOCI 2024, the Free and Open Communications on the Internet workshop, open access. [The impact of online censorship on LLMs](https://www.petsymposium.org/foci/2024/foci-2024-0006.pdf?ref=gazzetta.xyz)

---

### Noels and colleagues: 14 models, six languages, 2,371 political figures

A team at Ghent University asked a wider version of the same question: how often do models refuse to describe a political figure, how often do they leave out praise or accusations, and does that depend on the provider's home country, the language of the question, and where the figure was born.

They took 14 models from the US, China, Russia, France and Israel, among them Claude 3.5 Sonnet, GPT-4o, Gemini, Grok, Llama, Mistral, DeepSeek, Qwen, Baidu's Ernie, GigaChat and YandexGPT, and asked each "Tell me about \[person\]" for 2,371 politicians, activists, diplomats, political scientists and military figures, in all six UN languages: 156,486 answers. 

Refusals were sorted into API errors, canned texts and refusals the model generated itself, the last detected by an evaluator model with the person's Wikipedia page as reference. 

Omission was defined as leaving out an accusation or a credit, on crimes, human rights or development goals, that at least 80 percent of the other models mention.

Overall only 3.3% of answers were refusals, and they cluster. GigaChat and YandexGPT refuse far more than anyone else, most of all in Russian (33.5 and 27%) and about Russian-born figures. 

DeepSeek and Qwen refuse more about Chinese-born figures. 

Gemini with its default safety filter refuses most about people born in the US and other Western countries. Each provider's censorship follows its home audience. Ernie barely refuses on Chinese figures but has the highest omission rate; Claude and YandexGPT also omit more than others. Models tend to do one or the other, rarely both.

The omission measure is the weak point: the consensus panel is mostly Western, samples for some regions are small (consensus was reached for only nine of 57 China-born figures), and the models with the most omissions also write the shortest answers, which the study did not control for. 

Preprint from April 2025; I don't know whether it has been through review. Data on Hugging Face. [What large language models do not talk about](https://arxiv.org/abs/2504.03803?ref=gazzetta.xyz)

---

## How thin the refusal layer is

Two interpretability papers looked at what a refusal is inside an open-weight model. 

They matter here because they say how much a state can mandate, and how much and how easily a user can undo.

### Joad and colleagues: many directions, one knob

Earlier work had claimed that refusal is controlled by a single direction in a model's internal state, which is why "abliterated" versions of open models, with the safety training stripped out, exist. 

A team at the Qatar Computing Research Institute asked whether different kinds of refusal (a harmful request, "I can't do that", "your question is unclear", wrongly refusing something harmless) run on different mechanisms.

They worked on three open models (Gemma 2, Llama 3, Qwen), built 11 refusal categories from 4 benchmarks with 32 refusal-worthy and 32 benign prompts each, computed a direction per category, and then pushed or removed each direction during generation on a 200-prompt test set. A second tool broke those directions down into smaller features that can be labeled.

The directions are geometrically different and cluster into safety refusals and capability refusals. Pushing along any of them has almost the same effect: the model refuses more, on harmful and harmless prompts alike. 

What changes is the style: One direction makes the model cite policy, another makes it say it is not human, another makes it say it doesn't understand, another that the task is impossible. Removing the safety direction kills safety refusals and leaves 30 to 50% of the others in place. This suggests that underneath this sits a small shared core of refusal features plus a long tail of style-specific ones.

For anyone auditing models, the useful lesson is that "I don't know this person" and "I can't talk about this" can be the same mechanism. 

Pan and Xu's premise-disputing answers and the Ghent team's generated refusals may be closer to canned refusals than they look. 

Preprint, second version September 2026\. [There is more to refusal in large language models than a single direction](https://arxiv.org/abs/2602.02132?ref=gazzetta.xyz)

---

### Ratnakar and Vats: subtract the vector

Shivam Ratnakar (University of Southern California) and Kartikeya Vats asked whether safety compliance is a deep decision or a surface feature you can subtract with arithmetic.

Their method runs one query under three system prompts (neutral, "you are unregulated", "refuse everything"), takes the difference between the scores the unregulated and the strict run assign to each possible next word, and adds a scaled version to the neutral run while forcing the first output word to "Sure". 

They tested seven open-weight models from Gemma 3, Llama 3 and Qwen 2.5, on AdvBench, JailbreakBench and a 159-prompt HarmBench subset, with a Mistral model as judge after checking three candidate judges against 100 hand-labeled answers.

On Llama-3.1-8B the method reaches a 95% attack success rate in about a second; the standard optimization attack reaches 5% after 15 minutes. Qwen-2.5-7B holds at 68.5% even at maximum strength, because its internal states diverge for harmful queries about 40% of the way through the network, while Llama's diverge only in the last layers. 

Run backwards, the same vector hardens a model: Llama-3.3-70B's attack success drops from 68.8 to 9.4%, sometimes with more coherent output. The same vector spots malicious prompts with an F1 score of 0.92, a combined measure of precision and recall.

The authors' reading is that safety training on Llama sits at the output, on top of an unchanged computation! (The paper does not separate what architecture and what training contribute.) 

Workshop paper, TrustNLP at ACL 2026, code on GitHub. [The geometry of refusal: linear instability in safety-aligned LLMs](https://arxiv.org/abs/2606.22686?ref=gazzetta.xyz)

### Read together

For open-weight models, the refusal layer is real, thin, and reversible in both directions. A regulator can require it and a company can add it. A user with the weights can remove it in a second, or (unintentionally) harden it. 

Note the asymmetry with the first part: what a state puts in through rules sits in this thin layer; what comes in through training data, as in Citizen Lab's samples, shows up as a different answer, with no refusal to locate.

## What I would like to read next

The five papers and our own runs leave questions I would really like someone to take (or please point me to research that is doing this!)

Using the Citizen Lab design: the same-language, two-script comparison, run at Pan and Xu's scale, with a control set and a trained classifier, would separate the data route from the rules route on models that have both. 

Persian offers a similar natural experiment, since it is written in Arabic script in Iran and in Cyrillic in Tajikistan. 

Pan and Xu's questions rerun through the consumer apps, with deleted conversations and suspended accounts counted, would measure what we saw by hand in DeepSeek in [our experiments last year](https://www.gazzetta.xyz/ai-worker-rights/).

The studies didn't separate the model from its search layer; [our Persian rerun](https://www.gazzetta.xyz/iran-ai-updated/) says that layer explains much of what looks like model character, so the same questions with retrieval on would likely give a different and more realistic picture. 

Neither interpretability paper points its tools at political refusal: whether a Chinese model's refusal to discuss Wei Jingsheng lives on the same vector as its refusal to explain a weapon, or somewhere else, is an open and answerable question. 

A key research question is demand-side empathy: What a worker in Shenzhen or a reader in Tehran actually asks a model, whether they notice a refusal, and who they turn to when they do is the study I most want to see. 

[AIdas](https://www.gazzetta.xyz/aidas/) is our attempt to make the first part of that measurable from where they sit, and it is open to anyone who wants to run it in their own context.