{"rewrite":{"id":"r_bb7eacacfbd02e69b8aa1432","clusterId":"c_ef2e72074ae31d165fb62684","slug":"z-ai-details-glm-5-3-flash-service-built-on-100-000-chinese-ai-accelerators","model":"deepseek-v4-1-flash","headline":"Z.ai Details GLM-5.3-Flash Service Built On 100,000 Chinese AI Accelerators","summary":"Z.ai explained in a blog post how it built the inference infrastructure behind the production service for GLM-5.3-Flash, running on a cluster of more than 100,000 Chinese-made AI accelerators. The company says most of the work was carried out by an infrastructure agent powered by GLM-5.3, and that optimizations brought end-to-end service performance to about 3x and cost per token in line with mainstream NVIDIA GPUs.","whyItMatters":"The account is a rare public description of a production inference stack running at that cluster scale on Chinese-made accelerators, and it credits an agent-driven feedback loop, not an infrastructure team, with most of the work.","webCardHtml":"\u003cp\u003eGLM-5.3-Flash also ran under the name Ox-Alpha on OpenCode and OpenRouter. Z.ai says that within one week of launch it was the most-used AI model on both platforms, processing more than 62 trillion tokens in six days.\u003c/p\u003e\u003cp\u003eThe engineering account covers memory optimization that trades computation for bandwidth, intra-node tensor parallelism for linear attention and the LM head, ReplaySSM, W8A8 quantization, mixed-precision cache quantization using INT8, FP8 and BF16, layer splitting, and an encode-prefill-decode disaggregated architecture. From initial model adaptation to production operation took just under two weeks.\u003c/p\u003e","blueskyPost":"Z.ai says an agent running on GLM-5.3 did most of the work tuning its own serving cluster. That puts the model in the position of writing the optimizations that decide its own inference cost.","twitterPost":"Z.ai says a GLM-5.3 agent did most of the work tuning the cluster serving GLM-5.3-Flash. The model is optimizing its own serving cost.","threadsPost":"Z.ai says most of the infrastructure work behind GLM-5.3-Flash was done by an agent running on GLM-5.3 itself. The model tuned the cluster that serves the model. Cost per token on those 100,000 accelerators now lands near mainstream NVIDIA GPUs.","newsletterBlurb":"Z.ai published an account of building the inference infrastructure behind GLM-5.3-Flash on a cluster of more than 100,000 Chinese-made AI accelerators. The company says an infrastructure agent powered by GLM-5.3 carried out most of the work, and that memory and architecture optimizations lifted end-to-end performance by about 3x.","attributionJson":"[{\"source\":\"GIGAZINE\",\"url\":\"https://gigazine.net/news/20260918-how-glm-built-inference-infrastructure/\",\"title\":\"Z.ai Shares Know-How from Running 'GLM-5.3-Flash' Production Service on Chinese AI Infrastructure, Managing Infrastructure with AI Agents to Match NVIDIA GPU Efficiency\"}]","lintFlagsJson":null,"lintHits":0,"costUsd":0,"inputTokens":4637,"outputTokens":660,"status":"published","repairAttempts":0,"nextRepairAt":null,"factsAttemptedAt":1789709372,"createdAt":"2026-09-18T05:25:20.000Z","publishedAt":"2026-09-18T05:28:25.000Z","updatedAt":"2026-09-18T05:28:25.000Z"},"cluster":{"id":"c_ef2e72074ae31d165fb62684","canonicalTitle":"Z.aiが中国製AIインフラで「GLM-5.3-Flash」の本番サービスを提供したノウハウを共有、AIエージェントでインフラを管理してNVIDIA GPUと同等まで効率化","representativeArticleId":"a_63fef05042161671665de841","sourceCount":1,"writtenSourceCount":1,"writeAttempts":0,"isSolo":true,"entitiesJson":"{\"anime_titles\":[],\"manga_titles\":[],\"work_titles\":[\"GLM-5.3-Flash\"],\"studios\":[\"Z.ai\"],\"people\":[],\"type\":\"news\",\"domain\":\"other\",\"is_roundup\":false}","contentType":"news","status":"published","firstSeenAt":"2026-09-18T04:29:00.000Z","lastSeenAt":"2026-09-18T04:29:00.000Z","updatedAt":"2026-09-18T05:28:21.000Z"},"attribution":[{"source":"GIGAZINE","url":"https://gigazine.net/news/20260918-how-glm-built-inference-infrastructure/","title":"Z.aiが中国製AIインフラで「GLM-5.3-Flash」の本番サービスを提供したノウハウを共有、AIエージェントでインフラを管理してNVIDIA GPUと同等まで効率化"}],"entities":{"anime_titles":[],"manga_titles":[],"work_titles":["GLM-5.3-Flash"],"studios":["Z.ai"],"people":[],"type":"news","domain":"other","is_roundup":false},"keyFacts":null}
