{"rewrite":{"id":"r_8ad758a7cb66fa91f8550492","clusterId":"c_4632a56d0ff3e0544b78b7b5","slug":"strata-runs-a-125b-model-on-gaming-pcs","model":"deepseek-v4-1-flash","headline":"Strata Runs A 125B Model On Gaming Pcs","summary":"Strata is an app that runs a quantized version of the 125-billion-parameter Qwen3.8-Flash-Next on consumer gaming PCs. The repository lists requirements of an NVIDIA or AMD GPU with 12GB or more of VRAM, 32GB or more of RAM, 80GB or more of storage, and Windows 10/11 or Linux. Reported runs hit 94 tokens/second on an RTX 5070 with 64GB RAM, but the 2-bit quantized build produced looping output in one prompt.","whyItMatters":"The reported token speeds are real, but the looping pi output shows what the quantized tiers cost: Strata makes a 125B model runnable on a desktop, not fully intact.","webCardHtml":"\u003cp\u003eThe repo splits the model by memory. With 32GB of RAM only the GSQ-RCO Coder build runs, a lightweight version with non-coding expert models removed. At 64GB the 2-bit IQ2_XS and the 3-bit IQ3_XXS and IQ3_S builds fit, and higher-precision quantizations need more still.\u003c/p\u003e\u003cp\u003eSpeed reports circulated on X. One post measured Q2_0 at 94 tokens/second on an RTX 5070, Ryzen 5 7600 and 64GB RAM, and IQ3_S at 53 tokens/second. Another measured Q2_0 at 60 tokens/second on a Radeon RX 9070 XT, Ryzen 9 3900X and 47GB RAM. A third reported 140 tokens/second for IQ2_XS on a 4090, then showed the same build looping on a pi prompt.\u003c/p\u003e","blueskyPost":"Strata's 2-bit quantized Qwen3.8-Flash-Next hit 94 tokens/second on an RTX 5070, but the same build looped on one prompt. The 2-bit compression is where the speed comes from and where the output breaks.","twitterPost":"Strata hit 94 tokens/second on an RTX 5070. The same 2-bit quantized Qwen3.8-Flash-Next build looped on one prompt.","threadsPost":"Strata's speed number and its failure number come from the same place. The 2-bit quantized Qwen3.8-Flash-Next ran at 94 tokens/second on an RTX 5070 with 64GB RAM, and in one prompt that same build produced looping output. The compression is the trade.","newsletterBlurb":"Strata is an app for running a quantized version of the 125-billion-parameter Qwen3.8-Flash-Next on consumer gaming PCs. Users reported 94 tokens/second on an RTX 5070 with 64GB RAM and 60 tokens/second on a Radeon RX 9070 XT, while a 2-bit build looped on a pi prompt. The repository lists 12GB of VRAM, 32GB of RAM and 80GB of storage as the floor.","attributionJson":"[{\"source\":\"GIGAZINE\",\"url\":\"https://gigazine.net/news/20261005-strata-local-ai/\",\"title\":\"Strata: An AI Execution App That Enables High-Speed Local Execution of a Quantized Version of the 125B 'Qwen3.8-Flash-Next' on a Gaming PC\"}]","lintFlagsJson":null,"lintHits":0,"costUsd":0,"inputTokens":4900,"outputTokens":756,"status":"published","repairAttempts":0,"nextRepairAt":null,"factsAttemptedAt":1791268027,"createdAt":"2026-10-06T06:21:13.000Z","publishedAt":"2026-10-06T06:26:05.000Z","updatedAt":"2026-10-06T06:26:05.000Z"},"cluster":{"id":"c_4632a56d0ff3e0544b78b7b5","canonicalTitle":"125Bの「Qwen3.8-Flash-Next」の量子化版をゲーミングPCで高速ローカル実行可能にするAI実行アプリ「Strata」","representativeArticleId":"a_18d478c058baa89055126b83","sourceCount":1,"writtenSourceCount":1,"writeAttempts":0,"isSolo":true,"entitiesJson":"{\"anime_titles\":[],\"manga_titles\":[],\"work_titles\":[\"Qwen3.8-Flash-Next\",\"Strata\"],\"studios\":[],\"people\":[],\"type\":\"news\",\"domain\":\"other\",\"is_roundup\":false}","contentType":"news","status":"published","firstSeenAt":"2026-10-05T12:08:00.000Z","lastSeenAt":"2026-10-05T12:08:00.000Z","updatedAt":"2026-10-06T06:26:06.000Z"},"attribution":[{"source":"GIGAZINE","url":"https://gigazine.net/news/20261005-strata-local-ai/","title":"125Bの「Qwen3.8-Flash-Next」の量子化版をゲーミングPCで高速ローカル実行可能にするAI実行アプリ「Strata」"}],"entities":{"anime_titles":[],"manga_titles":[],"work_titles":["Qwen3.8-Flash-Next","Strata"],"studios":[],"people":[],"type":"news","domain":"other","is_roundup":false},"keyFacts":null}
