{"rewrite":{"id":"r_b8c2d0190efd585c6ca2ff4f","clusterId":"c_2e293582d33dd3895aa829d9","slug":"google-unveils-gemini-3-5-transcribe-speech-model","model":"deepseek-v4-flash","headline":"Google Unveils Gemini 3.5 Transcribe Speech Model","summary":"Google announced Gemini 3.5 Transcribe, a speech recognition model that converts raw audio into clean, formatted text in real time while automatically removing filler words like \"um\" and \"uh.\" The model also reflects speaker self-corrections and applies formatting such as bullet points and parentheses, turning rambling speech into polished prose. It is available through two APIs: the Live API for continuous bidirectional streaming with sub-second latency, and the Interactions API for recorded audio with speaker identification and word-level timestamps. Google reports a word error rate of 4.0% for streaming and 2.6% for non-streaming, as measured by Artificial Analysis, and a 70% reduction in time to final transcription compared to the earlier Chirp 3 model. The model supports over 85 languages and allows custom vocabulary for technical terms. Gemini 3.5 Transcribe is already integrated into Android's Rambler voice input and the macOS Gemini app's Speak to Window feature. Developers can access it through Google AI Studio, Google Antigravity, and the Gemini Enterprise Agent Platform. Pricing is estimated at about $0.005 per minute for recorded audio and $0.009 per minute for real-time, with free tiers available.","whyItMatters":"The announcement positions Gemini 3.5 Transcribe against a competitive field where its non-streaming error rate of 2.6% ranks fifth on Artificial Analysis's leaderboard, behind models like ElevenLabs' Scribe v2 and Microsoft's MAI-Transcribe-1.5, making it a benchmark comparison rather than a clear top performer.","webCardHtml":"\u003cp\u003e\u003c/p\u003e\u003cul\u003e\n\u003cli\u003eGIGAZINE frames the model for voice agents, real-time captioning tools, and call analysis pipelines, and notes it \u0026#34;represents a significant advancement\u0026#34; over Chirp 3 in Artificial Analysis measurements.\u003c/li\u003e\n\u003cli\u003eGameBusiness.jp details a side-by-side test where the conventional voice input misheard phone names as \u0026#34;Google Pixelmator\u0026#34; and \u0026#34;iPhone 7 Pro Max,\u0026#34; while Gemini 3.5 Transcribe read context correctly.\u003c/li\u003e\n\u003cli\u003eThe same test shows the model keeps a filler word when it is contextually necessary, rendering \u0026#34;いわゆるフィラー、あーとかうーとかが\u0026#34; as \u0026#34;いわゆるフィラー(あー、うー)が\u0026#34; rather than deleting it.\u003c/li\u003e\n\u003cli\u003eSpeaker separation supports up to 3 speakers reliably, with 4 or more experimental; the API documentation allows up to 8 speakers.\u003c/li\u003e\n\u003cli\u003eAudio length is capped at 1 hour, reduced to 30 minutes when speaker separation or timestamps are enabled.\u003c/li\u003e\n\u003cli\u003eDevelopers can register up to 1,000 custom vocabulary terms.\u003c/li\u003e\n\u003cli\u003eBeyond transcription, the model supports function calling, such as having another model generate images based on content.\u003c/li\u003e\n\u003cli\u003eGboard\u0026#39;s AI Voice Input (Rambler) is a key feature of the Pixel 11 series and also supports editing specific places by voice or rewriting the style of entire text after input.\u003c/li\u003e\n\u003cli\u003emacOS Gemini app access uses the Fn/globe key as a shortcut by default, working in text input areas of any app or window.\u003c/li\u003e\n\u003cli\u003eAt the May 2026 Gemini Intelligence announcement, AI voice input was cited as a feature with rollout promised to the latest Galaxy devices.\u003c/li\u003e\n\u003cli\u003eChrome and Gemini Enterprise for Customer Experience are listed as \u0026#34;coming soon\u0026#34; with timing undecided.\u003c/li\u003e\n\u003c/ul\u003e\u003cp\u003e\u003c/p\u003e","blueskyPost":"Gemini 3.5 Transcribe's 70% faster transcription than Chirp 3 suggests Google is prioritizing speed over raw accuracy, with the streaming error rate at 4.0%.","twitterPost":"Gemini 3.5 Transcribe's 70% speed gain over Chirp 3 trades off accuracy: streaming error is 4.0%.","threadsPost":"Gemini 3.5 Transcribe's 70% faster transcription than Chirp 3 comes with a tradeoff: streaming error is 4.0%, higher than non-streaming's 2.6%. Google is betting real-time editing matters more than perfect recognition.","newsletterBlurb":"Google unveiled Gemini 3.5 Transcribe, a speech model that cleans up filler words and self-corrections in real time. It's already in Pixel 11 and macOS Gemini app voice input, with APIs for developers starting at about $0.005 per minute.","attributionJson":"[{\"source\":\"GIGAZINE\",\"url\":\"https://gigazine.net/news/20260827-gemini-3-5-transcribe/\",\"title\":\"Google announces speech recognition model 'Gemini 3.5 Transcribe' that automatically removes 'um' and 'uh' while transcribing in real time\"},{\"source\":\"GameBusiness.jp\",\"url\":\"https://www.gamebusiness.jp/article/2026/08/28/27808.html\",\"title\":\"Gemini 3.5 Transcribe: High-Performance AI Transcription That Cleans Up Even Thinking-Out-Loud Speech, Reflects Corrections, Offered in Gboard and Mac Gemini App\"}]","lintFlagsJson":null,"lintHits":0,"costUsd":0,"inputTokens":12052,"outputTokens":1594,"status":"published","repairAttempts":0,"nextRepairAt":null,"factsAttemptedAt":1787891406,"createdAt":"2026-08-28T04:22:07.000Z","publishedAt":"2026-08-28T04:26:45.000Z","updatedAt":"2026-08-28T04:22:07.000Z"},"cluster":{"id":"c_2e293582d33dd3895aa829d9","canonicalTitle":"「あー」「えー」などを自動削除しつつリアルタイムで文字起こししできる音声認識モデル「Gemini 3.5 Transcribe」をGoogleが発表","representativeArticleId":"a_e21f9cb91d2b1e584b8c1216","sourceCount":2,"writtenSourceCount":2,"writeAttempts":0,"isSolo":false,"entitiesJson":"{\"anime_titles\":[],\"manga_titles\":[],\"work_titles\":[\"Gemini 3.5 Transcribe\"],\"studios\":[],\"people\":[],\"type\":\"announcement\",\"domain\":\"other\",\"is_roundup\":false}","contentType":"news","status":"published","firstSeenAt":"2026-08-27T04:15:00.000Z","lastSeenAt":"2026-08-27T22:00:03.000Z","updatedAt":"2026-08-28T04:26:46.000Z"},"attribution":[{"source":"GameBusiness.jp","url":"https://www.gamebusiness.jp/article/2026/08/28/27808.html","title":"考えながら話しても清書レベルの高性能AI文字起こしGemini 3.5 Transcribe―言い直しも反映、GboardやMac版Geminiアプリに提供"},{"source":"GIGAZINE","url":"https://gigazine.net/news/20260827-gemini-3-5-transcribe/","title":"「あー」「えー」などを自動削除しつつリアルタイムで文字起こししできる音声認識モデル「Gemini 3.5 Transcribe」をGoogleが発表"}],"entities":{"anime_titles":[],"manga_titles":[],"work_titles":["Gemini 3.5 Transcribe"],"studios":[],"people":[],"type":"announcement","domain":"other","is_roundup":false},"keyFacts":["Google announced Gemini 3.5 Transcribe on August 26, 2026 (early August 27 Japan time).","Gemini 3.5 Transcribe achieves a word error rate of 4.0% for streaming and 2.6% for non-streaming, as measured by Artificial Analysis.","The model supports over 85 languages and allows custom vocabulary with up to 1,000 custom terms.","Pricing is estimated at about $0.005 per minute for recorded audio and $0.009 per minute for real-time transcription, with free tiers.","Gemini 3.5 Transcribe is already available in Gboard's AI Voice Input on Pixel 11 series devices and the macOS Gemini app's Speak to Window feature."]}
