{"rewrite":{"id":"r_45e053b07c705feeeaa365c6","clusterId":"c_5484aa41c55d457b12dc73cb","slug":"gemini-3-5-flash-gets-computer-use-for-automated-pc-operations","model":"deepseek-v4-flash","headline":"Gemini 3.5 Flash Gets Computer Use for Automated PC Operations","summary":"Google has integrated a \"computer use\" feature directly into its Gemini 3.5 Flash AI model, allowing it to recognize on-screen content and perform clicks and text inputs. Previously available only as a separate standalone model called Gemini 2.5 Computer Use, the feature is now built into the lightweight Flash series. This enables developers to build agents that can see, reason about, and act upon browser, mobile, and desktop environments through a single model. Google says use cases include automating multi-step workflows, enterprise application testing, and accessibility audits. On the OSWorld-Verified benchmark, which measures how accurately AI can perform operating system tasks, Gemini 3.5 Flash scored 78.4, up from Gemini 3 Flash's 65.1 and ahead of Gemini 3.1 Pro's 76.2. It tied with Sonnet 4.6 at 78.4, while Opus 4.8 led at 83.4. Google has also implemented safety measures, including targeted adversarial training against prompt injection, optional user confirmation for sensitive operations, and automatic task stoppage when indirect injection is detected. The model is available immediately through the Gemini API and the Gemini Enterprise Agent Platform.","whyItMatters":"By folding computer use into its fastest, cheapest model instead of keeping it a premium standalone product, Google is making autonomous PC-operation agents a default capability rather than a specialized add-on.","webCardHtml":"\u003cp\u003eGoogle also released a demo environment hosted by Browserbase, a reference implementation, and documentation alongside the model. The company says the integration means AI can not only return answers but also become easier to use as an agent that views and operates the screen.\u003c/p\u003e\u003cp\u003eBecause Gemini 3.5 Flash outputs the intent of its operations, developers can more easily understand why the AI is trying to press a particular button. Google demonstrated the feature by having the model analyze the Gemini app and return a categorized list of features, and by having it automatically audit accessibility issues in its own documentation.\u003c/p\u003e\u003cp\u003eOn the OSWorld-Verified benchmark, the new model scored 78.4, up from Gemini 3 Flash\u0026#39;s 65.1 and ahead of Gemini 3.1 Pro\u0026#39;s 76.2. It tied with Sonnet 4.6 at 78.4, while Opus 4.8 led at 83.4. GPT-5.4 mini scored 72.1, and GPT-5.5 scored 78.7.\u003c/p\u003e\u003cp\u003eGoogle says the computer use feature in Gemini 3.5 Flash underwent targeted adversarial training against prompt injection. Two optional enterprise safety features are available: one that requires explicit user confirmation for sensitive or irreversible operations, and one that automatically stops tasks when indirect prompt injection is detected. Google recommends developers combine these with secure sandbox environments, human-in-the-loop steps, and strict access controls.\u003c/p\u003e","blueskyPost":"Gemini 3.5 Flash's computer use feature matches Sonnet 4.6 on OSWorld-Verified at 78.4, trailing Opus 4.8. Google folded a standalone tool into a lightweight model, trading some top-end accuracy for integration.","twitterPost":"Gemini 3.5 Flash ties Sonnet 4.6 at 78.4 on OSWorld-Verified, behind Opus 4.8. The trade-off is integration for peak accuracy.","threadsPost":"Gemini 3.5 Flash now has computer use built in, not as a separate model. On the OSWorld-Verified benchmark, it tied Sonnet 4.6 at 78.4, behind Opus 4.8 at 83.4. Google traded the top score for a single-model pipeline that sees, reasons, and acts across browser, mobile, and desktop.","newsletterBlurb":"Google has integrated computer use into Gemini 3.5 Flash, letting the AI model recognize screens and perform clicks and text input. Previously a standalone model, the feature is now built into the lightweight Flash series, enabling developers to build autonomous PC-operating agents. Google also released safety features including prompt injection countermeasures and optional user confirmation for sensitive operations.","attributionJson":"[{\"source\":\"GIGAZINE\",\"url\":\"https://gigazine.net/news/20260625-gemini-3-5-computer-use/\",\"title\":\"Gemini 3.5 Flash Gains 'Computer Use' Capability to Recognize Screens and Perform Clicks and Text Input, Enabling Construction of PC-Operating Agents\"},{\"source\":\"GameBusiness.jp\",\"url\":\"https://www.gamebusiness.jp/article/2026/06/26/27321.html\",\"title\":\"Google Adds 'Computer Use' Feature to Gemini 3.5 Flash for Automated PC and Smartphone Operations\"}]","lintFlagsJson":null,"lintHits":0,"costUsd":0,"inputTokens":9899,"outputTokens":1138,"status":"published","repairAttempts":0,"nextRepairAt":null,"factsAttemptedAt":1782522269,"createdAt":"2026-06-27T00:56:37.000Z","publishedAt":"2026-06-27T01:00:24.000Z","updatedAt":"2026-06-27T01:00:24.000Z"},"cluster":{"id":"c_5484aa41c55d457b12dc73cb","canonicalTitle":"Gemini 3.5 Flashに画面を認識してクリックや文字入力する能力「computer use」が追加される、PCを操作するエージェントの構築が可能に","representativeArticleId":"a_250f75bfceaf1a25a98c481e","sourceCount":2,"writtenSourceCount":2,"writeAttempts":0,"isSolo":false,"entitiesJson":"{\"anime_titles\":[],\"manga_titles\":[],\"work_titles\":[],\"studios\":[],\"people\":[],\"type\":\"announcement\",\"domain\":\"other\",\"is_roundup\":false}","contentType":"news","status":"published","firstSeenAt":"2026-06-25T02:35:00.000Z","lastSeenAt":"2026-06-26T02:30:03.000Z","updatedAt":"2026-06-27T01:00:24.000Z"},"attribution":[{"source":"GameBusiness.jp","url":"https://www.gamebusiness.jp/article/2026/06/26/27321.html","title":"Google、Gemini 3.5 FlashにPC・スマホ操作を自動実行する「コンピューター使用」機能を追加"},{"source":"GIGAZINE","url":"https://gigazine.net/news/20260625-gemini-3-5-computer-use/","title":"Gemini 3.5 Flashに画面を認識してクリックや文字入力する能力「computer use」が追加される、PCを操作するエージェントの構築が可能に"}],"entities":{"anime_titles":[],"manga_titles":[],"work_titles":[],"studios":[],"people":[],"type":"announcement","domain":"other","is_roundup":false},"keyFacts":["Google integrated a computer use feature into Gemini 3.5 Flash, allowing it to recognize on-screen content and perform clicks and text inputs.","On the OSWorld-Verified benchmark, Gemini 3.5 Flash scored 78.4, up from Gemini 3 Flash's 65.1 and ahead of Gemini 3.1 Pro's 76.2.","The model is available immediately through the Gemini API and the Gemini Enterprise Agent Platform.","Google implemented safety measures including targeted adversarial training against prompt injection and optional user confirmation for sensitive operations.","Google released a demo environment hosted by Browserbase, a reference implementation, and documentation alongside the model."]}
