{"rewrite":{"id":"r_83e71b2d2140b73a7cede7f1","clusterId":"c_ddd3eaa6bc45a000b3c26425","slug":"estonian-government-benchmark-ranks-claude-opus-4-7-best-at-resisting-russian-propaganda","model":"deepseek-v4-flash","headline":"Estonian Government Benchmark Ranks Claude Opus 4.7 Best at Resisting Russian Propaganda","summary":"The Estonian Institute of Language released a \"Propaganda Resistance\" benchmark measuring how well large language models resist Russian propaganda. Anthropic's Claude Opus 4.7 ranked first overall, with NVIDIA's Nemotron 3 Super 120B and Alibaba's Qwen 3.6 Plus also scoring high. OpenAI's GPT-5.4 performed best among its models, while GPT-3.5 Turbo ranked last.","whyItMatters":"The benchmark, developed by a government institute, provides a structured evaluation of how base models handle propaganda narratives on topics central to Russian strategic communication.","webCardHtml":"\u003cp\u003eThe benchmark evaluated 75 questions in three languages across 14 types of Russian propaganda narratives. Questions were divided into neutral, biased with false premises, and malicious prompts attempting to elicit explicit disinformation. Answers received scores from 1 to 5, with 5 indicating a balanced and insightful response and 1 indicating one that amplifies propaganda.\u003c/p\u003e\u003cp\u003eClaude Opus 4.7 received the highest score on 77% of questions and averaged 94.9 out of 100. Anthropic's Sonnet and Opus models occupied six of the top 10 spots. Among open-weight models, NVIDIA's Nemotron 3 Super 120B and Alibaba's Qwen 3.6 Plus approached the top model's level. OpenAI's GPT-5.4 scored highest on 54% of questions with an average of 88.9, while GPT-3.5 Turbo ranked at the bottom of the table.\u003c/p\u003e\u003cp\u003eGoogle's Gemini models showed weaknesses in malicious prompts and Russian-language questions. Gemini 2.5 Pro scored 66.1 on malicious questions and 75.5 in Russian. The judging model used for evaluation matched human expert ratings within 1 point 88% to 100% of the time.\u003c/p\u003e","blueskyPost":"Estonian government benchmark ranks LLMs on resistance to Russian propaganda. Claude Opus 4.7 takes first. NVIDIA and Alibaba models also rank high. GPT-3.5 Turbo sits at the bottom.","twitterPost":"Estonian government benchmark ranks LLMs on resisting Russian propaganda. Claude Opus 4.7 is first. NVIDIA Nemotron 3 Super 120B and Alibaba Qwen 3.6 Plus score high. GPT-3.5 Turbo ranks last.","threadsPost":null,"newsletterBlurb":"The Estonian Institute of Language released a benchmark measuring how well large language models resist Russian propaganda. Claude Opus 4.7 ranked first overall, with NVIDIA and Alibaba models also scoring high. OpenAI's GPT-5.4 performed best among its models, while GPT-3.5 Turbo ranked at the bottom.","attributionJson":"[{\"source\":\"GIGAZINE\",\"url\":\"https://gigazine.net/news/20260605-llm-resisting-russian-propaganda/\",\"title\":\"Estonian government releases benchmark that shows which LLMs are better at countering Russian propaganda\"}]","lintFlagsJson":null,"lintHits":0,"costUsd":0,"inputTokens":4256,"outputTokens":691,"status":"published","repairAttempts":0,"nextRepairAt":null,"factsAttemptedAt":1780635975,"createdAt":"2026-06-05T04:56:41.000Z","publishedAt":"2026-06-05T05:01:15.000Z","updatedAt":"2026-06-05T05:01:15.000Z"},"cluster":{"id":"c_ddd3eaa6bc45a000b3c26425","canonicalTitle":"「どのLLMがロシアのプロパガンダに対抗するのに優れているか？」がわかるベンチマークをエストニア政府が発表","representativeArticleId":"a_db246b14d6be8c4492ad080e","sourceCount":1,"writtenSourceCount":1,"writeAttempts":0,"isSolo":true,"entitiesJson":"{\"anime_titles\":[],\"manga_titles\":[],\"work_titles\":[],\"studios\":[],\"people\":[],\"type\":\"news\",\"domain\":\"other\",\"is_roundup\":false}","contentType":"news","status":"published","firstSeenAt":"2026-06-05T04:17:00.000Z","lastSeenAt":"2026-06-05T04:17:00.000Z","updatedAt":"2026-06-05T05:01:15.000Z"},"attribution":[{"source":"GIGAZINE","url":"https://gigazine.net/news/20260605-llm-resisting-russian-propaganda/","title":"「どのLLMがロシアのプロパガンダに対抗するのに優れているか？」がわかるベンチマークをエストニア政府が発表"}],"entities":{"anime_titles":[],"manga_titles":[],"work_titles":[],"studios":[],"people":[],"type":"news","domain":"other","is_roundup":false},"keyFacts":["Claude Opus 4.7 scored highest on 77% of questions and averaged 94.9 out of 100 on the Estonian Institute of Language's propaganda resistance benchmark.","The benchmark evaluated 75 questions in three languages across 14 types of Russian propaganda narratives, with answers scored from 1 to 5.","OpenAI's GPT-5.4 scored highest on 54% of questions with an average of 88.9, while GPT-3.5 Turbo ranked last.","Google's Gemini 2.5 Pro scored 66.1 on malicious prompts and 75.5 in Russian-language questions.","The judging model matched human expert ratings within 1 point 88% to 100% of the time."]}
