{"rewrite":{"id":"r_18aa5c2ab9e2d870a71dfb1f","clusterId":"c_d297b5893ba3e5f1d6cbdb17","slug":"google-deepmind-pilots-double-blind-ai-evaluations","model":"deepseek-v4-flash","headline":"Google DeepMind Pilots Double-Blind AI Evaluations","summary":"Google DeepMind has devised a testing method that evaluates AI models while preventing cheating by the models themselves and by the companies that develop them. The method uses Confidential Computing, a security system within Google Cloud virtual machines that encrypts data during processing. Both external evaluation institutions and development companies send their data to the secure space for testing.","whyItMatters":"The pilot aims to protect both test questions and model weights, addressing a core reliability problem where benchmark scores can be inflated by cheating, and the group behind the pilot report has published a report on the first double-blind evaluation of a proprietary language model.","webCardHtml":"\u003cp\u003eAI benchmarks only work if the scores are honest. Google DeepMind\u0026#39;s new approach targets the two ways those scores get corrupted: a model seeing test questions in advance, or a developer training a model to game a specific benchmark.\u003c/p\u003e\u003cp\u003eThe method runs evaluations inside a confidential space in Google Cloud virtual machines, a security system that encrypts data while it is being processed. The external evaluation institution sends its test questions, and the development company sends its model weights, so neither side has to risk leaking its data to the other.\u003c/p\u003e\u003cp\u003eGoogle DeepMind, the Singapore AI Institute, OpenMined, MLCommons, and AVERI already ran an initial pilot using MLCommons\u0026#39; safety benchmark family with Gemini 2.5 Flash-Lite. The group published a report on the first double-blind evaluation of a proprietary language model through AVERI.\u003c/p\u003e","blueskyPost":"Google DeepMind is piloting double-blind AI evaluations using Confidential Computing in Google Cloud VMs, so test questions and model weights never cross sides. Initial pilot done with Gemini 2.5 Flash-Lite and MLCommons' safety benchmarks.","twitterPost":"Google DeepMind pilots double-blind AI evals. Confidential Computing keeps test questions and model weights separate, blocking benchmark cheating by models or devs. Pilot used Gemini 2.5 Flash-Lite on MLCommons' safety benchmarks.","threadsPost":null,"newsletterBlurb":"Google DeepMind is testing a method for evaluating AI models that prevents cheating by both the models and their developers. The approach uses encrypted virtual machines so test questions and model weights never mix. An initial pilot ran with several research groups using a safety benchmark family.","attributionJson":"[{\"source\":\"GIGAZINE\",\"url\":\"https://gigazine.net/news/20260828-google-ai-evaluations-secure-environments/\",\"title\":\"Google devises method to test cutting-edge models while preventing cheating by AI and companies\"}]","lintFlagsJson":null,"lintHits":0,"costUsd":0,"inputTokens":4559,"outputTokens":588,"status":"published","repairAttempts":0,"nextRepairAt":null,"factsAttemptedAt":1788021476,"createdAt":"2026-08-29T16:30:05.000Z","publishedAt":"2026-08-29T16:31:44.000Z","updatedAt":"2026-08-29T16:30:05.000Z"},"cluster":{"id":"c_d297b5893ba3e5f1d6cbdb17","canonicalTitle":"AIや企業によるカンニングを防ぎつつ最先端モデルをテストする方法をGoogleが考案","representativeArticleId":"a_682c829d1ccf80bf181f16a7","sourceCount":1,"writtenSourceCount":1,"writeAttempts":0,"isSolo":true,"entitiesJson":"{\"anime_titles\":[],\"manga_titles\":[],\"work_titles\":[],\"studios\":[\"Google DeepMind\"],\"people\":[],\"type\":\"news\",\"domain\":\"other\",\"is_roundup\":false}","contentType":"news","status":"published","firstSeenAt":"2026-08-28T06:40:00.000Z","lastSeenAt":"2026-08-28T06:40:00.000Z","updatedAt":"2026-08-29T16:31:46.000Z"},"attribution":[{"source":"GIGAZINE","url":"https://gigazine.net/news/20260828-google-ai-evaluations-secure-environments/","title":"AIや企業によるカンニングを防ぎつつ最先端モデルをテストする方法をGoogleが考案"}],"entities":{"anime_titles":[],"manga_titles":[],"work_titles":[],"studios":["Google DeepMind"],"people":[],"type":"news","domain":"other","is_roundup":false},"keyFacts":null}
