{"rewrite":{"id":"r_368d5b8fedfd110206ad8c75","clusterId":"c_13e04be2ec2a99e9949c7175","slug":"inception-ships-mercury-2-5-diffusion-llm-at-1107-tokens-per-second","model":"deepseek-v4-flash","headline":"Inception Ships Mercury 2.5 Diffusion LLM at 1107 Tokens Per Second","summary":"Inception announced Mercury 2.5, a diffusion-based language model that refines multiple tokens in parallel instead of generating text left to right. It outputs 1107 tokens per second, up from Mercury 2's 1009, with context length doubled to 260,000 tokens. Inception rates its quality on par with GPT-5.6 Luna (Low), Gemini 3.5 Flash-Lite, and Claude Haiku 4.5. Output pricing stays at $0.75 per million tokens while input pricing drops to $0.20, with an 80 percent launch discount on OpenRouter.","whyItMatters":"Mercury 2.5 puts diffusion-based generation on par with the cost-efficient frontier models GPT-5.6 Luna Low and Gemini 3.5 Flash-Lite while cutting input pricing, making parallel token generation a practical option for latency-sensitive agent workloads.","webCardHtml":"\u003cp\u003eMercury 2.5 applies the diffusion approach used in image generation to text, building a rough draft of the whole answer and refining multiple tokens in parallel instead of deciding one token at a time. Inception says the model matches the quality of GPT-5.6 Luna (Low), Gemini 3.5 Flash-Lite, and Claude Haiku 4.5.\u003c/p\u003e\u003cp\u003eThe company cites two deployments. AI phone agent developer OpenCall cut median response time to about 170 milliseconds, and coding AI firm Augment Code shortened context compression from about 150 seconds to 27 seconds with a 90 percent cost reduction.\u003c/p\u003e\u003cp\u003eInception also previewed Mercury Voice for speech AI and Mercury Router, which sends input to the appropriate model. The company says training on its largest next-generation model has begun, with a release targeted within the next few months.\u003c/p\u003e","blueskyPost":"Inception's Mercury 2.5 diffusion LLM hits 1107 tokens per second with a 260,000-token context. Inception rates it on par with GPT-5.6 Luna Low and Gemini 3.5 Flash-Lite. Input pricing drops to $0.20 per million tokens, with an 80 percent launch discount on OpenRouter.","twitterPost":"Mercury 2.5, Inception's diffusion-based LLM, outputs 1107 tokens per second with a 260,000-token context. Inception says quality matches GPT-5.6 Luna Low and Gemini 3.5 Flash-Lite. Input pricing falls to $0.20 per million tokens, with an 80 percent launch discount on OpenRouter.","threadsPost":null,"newsletterBlurb":"Inception announced Mercury 2.5, a diffusion-based language model that refines multiple tokens in parallel instead of generating left to right. It outputs 1107 tokens per second with a 260,000-token context, and Inception rates its quality on par with GPT-5.6 Luna Low and Gemini 3.5 Flash-Lite. Input pricing drops to $0.20 per million tokens, with an 80 percent launch discount on OpenRouter.","attributionJson":"[{\"source\":\"GIGAZINE\",\"url\":\"https://gigazine.net/news/20260909-mercury-2-5/\",\"title\":\"Mercury 2.5, an AI model that rapidly generates text via diffusion, arrives with 1107 tokens per second and GPT-5.6 Luna-class performance at about 120 yen per million output tokens\"}]","lintFlagsJson":null,"lintHits":0,"costUsd":0,"inputTokens":5852,"outputTokens":2613,"status":"published","repairAttempts":0,"nextRepairAt":null,"factsAttemptedAt":1788928168,"createdAt":"2026-09-09T04:25:25.000Z","publishedAt":"2026-09-09T04:28:20.000Z","updatedAt":"2026-09-09T04:28:20.000Z"},"cluster":{"id":"c_13e04be2ec2a99e9949c7175","canonicalTitle":"文章を「拡散」で高速生成するAIモデル「Mercury 2.5」登場、毎秒1107トークンの高速生成＆GPT-5.6 Luna級の性能を100万出力トークン当たり約120円で提供","representativeArticleId":"a_d53f97d186002b433033065d","sourceCount":1,"writtenSourceCount":1,"writeAttempts":0,"isSolo":true,"entitiesJson":"{\"anime_titles\":[],\"manga_titles\":[],\"work_titles\":[\"Mercury 2.5\"],\"studios\":[],\"people\":[],\"type\":\"announcement\",\"domain\":\"other\",\"is_roundup\":false}","contentType":"news","status":"published","firstSeenAt":"2026-09-09T03:30:00.000Z","lastSeenAt":"2026-09-09T03:30:00.000Z","updatedAt":"2026-09-09T04:28:20.000Z"},"attribution":[{"source":"GIGAZINE","url":"https://gigazine.net/news/20260909-mercury-2-5/","title":"文章を「拡散」で高速生成するAIモデル「Mercury 2.5」登場、毎秒1107トークンの高速生成＆GPT-5.6 Luna級の性能を100万出力トークン当たり約120円で提供"}],"entities":{"anime_titles":[],"manga_titles":[],"work_titles":["Mercury 2.5"],"studios":[],"people":[],"type":"announcement","domain":"other","is_roundup":false},"keyFacts":null}
