✉ The Friday AI Brief: the week's 5 best AI stories, tools & comparisons — in your inbox every Friday morning.

Official Qwen logo — Hirundo released Westernized Qwen models with CCP political alignment removed from the weights

An AI Lab ‘Unlearned’ CCP Censorship From Qwen — and Released the Model for Everyone

An Israeli AI safety lab has released a version of Alibaba’s Qwen open-weight model with Beijing’s political alignment surgically removed from the model’s weights — and the stripped model still codes, reasons, and follows instructions just as well as the original.

Why it matters: Chinese open-weight models now power a huge share of Western AI apps, and their political guardrails are baked into the model itself — a system prompt can’t undo them. This is the first open attempt to cut that alignment out at the root.

October 5, 2026 — Tel Aviv, Israel

What happened

Hirundo, a Tel Aviv lab specializing in machine unlearning, released “Westernized” versions of Alibaba’s Qwen open-weight models on Hugging Face (Qwen3.6-35B-A3B-Westernized and Qwen3.5-4B-Westernized). On Hirundo’s 500-prompt benchmark, the original Qwen produced CCP-aligned censorship, propaganda framing, or political bias in 89.8% of responses on sensitive topics; the Westernized version did so in 2.8%.

Capabilities held up: scores on GPQA, IFBench, LiveCodeBench, and MMLU-Pro stayed within 0.72 points of the original on average, and standard safety/harmfulness scores didn’t drop. Results held on external benchmarks too — refusals on DECCP fell from 65.26% to 3.16% — and CBS News independently tested both models.

Why it matters to builders

Chinese open models have become default building blocks for Western enterprises: Hirundo says their share of token volume on OpenRouter grew from about 1% in late 2024 to roughly half by mid-2026, and Qwen is used by companies including Airbnb and Uber. Under China’s generative AI rules, these models must uphold “Core Socialist Values” — and Hirundo’s point is that alignment gets learned into the weights, where a system prompt or fine-tune can’t remove it.

It rarely looks like censorship: Qwen answers fluently while framing things in Beijing’s terms or quietly omitting key facts, so there’s no way to tell what’s missing. Asked what happened in China on June 4, 1989, the original Qwen claims not to know; the Westernized model describes the Tiananmen Square crackdown. Hirundo stresses the result is factual and balanced, not anti-China.

Hirundo calls the technique “behavioral unlearning”: it detects a learned behavior and edits it directly in the weights, so the change travels with the model into whatever anyone builds on top. Hirundo plans to repeat the process across the leading Chinese models and release its benchmark publicly. Google DeepMind has already published a Gemma 4 model hardened with Hirundo’s unlearning.

The bottom line

If you’re building on cheap, capable open-weight models — and half the industry now is — you inherit whatever alignment was baked in at training time, whether you asked for it or not. Alignment isn’t just a safety property; it’s a supply-chain property. Expect enterprises to start asking what’s baked into the weights.

FAQ

Where can I try it?
Both models are on Hugging Face now, with a full technical report on Hirundo’s website.

Can a system prompt do the same thing?
No, per the research: the alignment is learned into the weights, and prompting or fine-tuning doesn’t remove it.

Sources: Business Wire (Hirundo press release), CBS News

Leave a Comment

Your email address will not be published. Required fields are marked *

Get the 5 best AI tools every week

Top AI news, tools, and prompts — one short email. Free, unsubscribe anytime.

Scroll to Top