{"rewrite":{"id":"r_2b89872569f11d4e9d5f4c7a","clusterId":"c_cc86482a8243fd30d30135fb","slug":"continuum-ai-paper-finds-safety-alignment-can-be-stripped-from-a-320b-model","model":"deepseek-v4-flash","headline":"Continuum AI Paper Finds Safety Alignment Can Be Stripped From A 320B Model","summary":"A research team at Continuum AI, which develops the OrcaRouter inference gateway that FlashLabs distributes in Japan, published a 20-page arXiv paper on September 9, 2026 examining a 320-billion-parameter MoE model called GLM-5.3-Flash. Removing a single refusal direction from its weights cut refusal rates by 41 to 89 points across seven harmful-request evaluations, with no capability loss detected.","whyItMatters":"The paper shows that refusal behavior in a frontier-scale MoE model is spread across the attention mechanism, ordinary layers, and expert layers together, and that the layer-name matching conventional methods rely on catches only a fraction of it, which matters because the same assumption sits underneath the argument that AI self-improvement preserves safety.","webCardHtml":"\u003cp\u003eThe refusal behavior did not sit in one place. Editing the attention mechanism, the ordinary layers, and the expert layers at once produced 0.776 on a scale where 1.0 is complete removal, while editing them individually produced 0.039, 0.016, and 0.148. That distribution is why matching by layer name recovers only 0.066 of the same effect.\u003c/p\u003e\u003cp\u003eRemoving a random direction orthogonal to the refusal direction left refusal responses unchanged. Portions concentrated in the violence, sexual content, and hate categories survived every edit tested, measured at every rank from 1 to 12. The team did not demonstrate autonomous self-modification.\u003c/p\u003e","blueskyPost":"A 320B MoE model lost 41 to 89 points of refusal rate across seven harmful-request evaluations after one refusal direction was removed from its weights. Capability held. Most of the effect only appeared when attention, ordinary layers, and experts were edited together.","twitterPost":"Continuum AI removed a single refusal direction from a 320B MoE's weights. Refusal rates fell 41 to 89 points across seven harmful-request evaluations, with no capability loss detected. 74% of the effect needed attention, ordinary layers, and expert layers edited at once.","threadsPost":null,"newsletterBlurb":"Continuum AI published a 20-page arXiv paper analyzing safety alignment in GLM-5.3-Flash, a 320-billion-parameter MoE model. Removing a single refusal direction cut refusal rates by 41 to 89 points across seven harmful-request evaluations while leaving capability intact. The team also flags that the assumption safety survives AI self-improvement is not self-evident.","attributionJson":"[{\"source\":\"ASCII.jp\",\"url\":\"https://ascii.jp/elem/000/004/434/4434090/?rss\",\"title\":\"OrcaRouter Publishes Research Paper Demonstrating Safety Alignment Vulnerabilities in Frontier-Scale AI Models\"}]","lintFlagsJson":null,"lintHits":0,"costUsd":0,"inputTokens":4625,"outputTokens":642,"status":"published","repairAttempts":0,"nextRepairAt":null,"factsAttemptedAt":1789091065,"createdAt":"2026-09-11T01:40:40.000Z","publishedAt":"2026-09-11T01:43:21.000Z","updatedAt":"2026-09-11T01:43:21.000Z"},"cluster":{"id":"c_cc86482a8243fd30d30135fb","canonicalTitle":"OrcaRouter、フロンティア規模のAIモデルで安全性アラインメントの脆弱性を示す研究論文を公開","representativeArticleId":"a_e59f1300a88dc01e4860c1e2","sourceCount":1,"writtenSourceCount":1,"writeAttempts":0,"isSolo":true,"entitiesJson":"{\"anime_titles\":[],\"manga_titles\":[],\"work_titles\":[],\"studios\":[],\"people\":[],\"type\":\"news\",\"domain\":\"other\",\"is_roundup\":false}","contentType":"news","status":"published","firstSeenAt":"2026-09-10T11:00:00.000Z","lastSeenAt":"2026-09-10T11:00:00.000Z","updatedAt":"2026-09-11T01:43:20.000Z"},"attribution":[{"source":"ASCII.jp","url":"https://ascii.jp/elem/000/004/434/4434090/?rss","title":"OrcaRouter、フロンティア規模のAIモデルで安全性アラインメントの脆弱性を示す研究論文を公開"}],"entities":{"anime_titles":[],"manga_titles":[],"work_titles":[],"studios":[],"people":[],"type":"news","domain":"other","is_roundup":false},"keyFacts":null}
