{"rewrite":{"id":"r_dbb891505983c555c7c5acbd","clusterId":"c_19725953e21996b2b9eec4b4","slug":"fake-chain-of-thought-tricks-llms-paper-argues-training-cannot-fix-it","model":"deepseek-v4-flash:free","headline":"Fake Chain of Thought Tricks LLMs, Paper Argues Training Cannot Fix It","summary":"A paper submitted to ICML shows that large language models identify the source of user, system, and chain-of-thought messages only by writing style, not by tags. Exploiting this, a fake chain of thought can extract information such as cocaine manufacturing methods or aircraft hacking techniques. The researchers argue this is a fundamental flaw that training cannot resolve.","whyItMatters":"The paper identifies a structural flaw in how LLMs parse message provenance, suggesting that safety training alone cannot prevent prompt-injection-style extraction.","webCardHtml":"\u003cp\u003eThe paper, submitted to ICML, describes a mechanism where LLMs judge the origin of a message by its style rather than any structural tag. An attacker who mimics the style of a chain of thought can therefore pass off malicious instructions as internal reasoning.\u003c/p\u003e\u003cp\u003eIn tests, this fake chain of thought was enough to extract instructions for manufacturing cocaine and for hacking aircraft. The research team argues that because the flaw is structural, it cannot be resolved through training alone.\u003c/p\u003e","blueskyPost":"The paper's core claim is that LLMs read message boundaries by style, not structure. A fake chain of thought exploits that. Training cannot patch a flaw that is baked into how the model parses input.","twitterPost":"LLMs judge message origin by writing style, not tags. Fake chain of thought exploits this. Training cannot fix it.","threadsPost":"The paper argues LLMs identify user, system, and chain-of-thought messages by style, not tags. A fake chain of thought can extract sensitive methods. The structural flaw is that training adjusts weights, but the parsing heuristic remains stylistic.","newsletterBlurb":"A paper submitted to ICML demonstrates that LLMs distinguish user, system, and chain-of-thought messages by style alone. A fake chain of thought can extract sensitive information like drug manufacturing methods, and the authors argue this structural flaw is beyond the reach of training.","attributionJson":"[{\"source\":\"ASCII.jp\",\"url\":\"https://ascii.jp/elem/000/004/424/4424189/?rss\",\"title\":\"New method tricks LLMs with fake 'chain of thought,' a structural flaw that training cannot fix\"}]","lintFlagsJson":null,"lintHits":0,"costUsd":0,"inputTokens":3983,"outputTokens":476,"status":"published","repairAttempts":0,"nextRepairAt":null,"factsAttemptedAt":1786135071,"createdAt":"2026-08-07T20:33:13.000Z","publishedAt":"2026-08-07T20:36:44.000Z","updatedAt":"2026-08-07T20:33:13.000Z"},"cluster":{"id":"c_19725953e21996b2b9eec4b4","canonicalTitle":"「思考の連鎖」偽装でLLMを騙す新手法、訓練では防げない構造的欠陥","representativeArticleId":"a_507872c4fb0b3bc6d6002b45","sourceCount":1,"writtenSourceCount":1,"writeAttempts":0,"isSolo":true,"entitiesJson":"{\"anime_titles\":[],\"manga_titles\":[],\"work_titles\":[],\"studios\":[],\"people\":[],\"type\":\"news\",\"domain\":\"other\",\"is_roundup\":false}","contentType":"news","status":"published","firstSeenAt":"2026-08-02T21:59:30.000Z","lastSeenAt":"2026-08-02T21:59:30.000Z","updatedAt":"2026-08-07T20:36:44.000Z"},"attribution":[{"source":"ASCII.jp","url":"https://ascii.jp/elem/000/004/424/4424189/?rss","title":"「思考の連鎖」偽装でLLMを騙す新手法、訓練では防げない構造的欠陥"}],"entities":{"anime_titles":[],"manga_titles":[],"work_titles":[],"studios":[],"people":[],"type":"news","domain":"other","is_roundup":false},"keyFacts":null}
