{"rewrite":{"id":"r_ef2d31f0a7bd456e4ce64476","clusterId":"c_e68ef8a1d06137189d00845f","slug":"flashlabs-ships-glm-5-3-mlx-runs-a-743b-model-on-a-512gb-mac","model":"deepseek-v4-flash","headline":"FlashLabs Ships GLM-5.3-MLX, Runs a 743B Model on a 512GB Mac","summary":"FlashLabs released GLM-5.3-MLX, an Apple Silicon-optimized MLX quantization of Z.ai's open-weight GLM-5.3 (about 743B total parameters, 40B active, 1M-token context), on Hugging Face. The OrcaSAQ method cuts the FP8 size from about 1.51TB to roughly 322GB at 2-bit, a 79 percent reduction, while keeping about 80.2 percent accuracy. A 512GB Mac with M3 Ultra runs the 2-bit, 3-bit, and 4-bit builds locally.","whyItMatters":"The full GLM-5.3, not a distilled or Flash variant, now runs locally on a single 512GB Mac, which is the step that moves a frontier open-weight model onto everyday hardware.","webCardHtml":"\u003cp\u003eThe release extends an August 28 build: FlashLabs already shipped GLM-5.3-Flash, a 320B-parameter variant, in MLX form, and GLM-5.3-MLX is the full model. OrcaSAQ needs no calibration data and keeps accuracy-critical tensors at higher bit widths, which is how the 2-bit build holds about 80.2 percent of FP8 accuracy while shrinking the footprint by roughly 79 percent.\u003c/p\u003e\u003cp\u003eThe four variants split by use case. A 512GB Mac with M3 Ultra runs the 2-bit and 3-bit builds comfortably and fits the 4-bit build. The 6-bit version, which reaches a cosine similarity of 0.9997 against FP8, targets two 512GB Macs linked through mlx.distributed or eight NVIDIA H200 GPUs. There is no 8-bit build because expert weights would grow larger than the FP8 source weights.\u003c/p\u003e\u003cp\u003eUnsloth also ships a GGUF version for llama.cpp environments. OrcaRouter\u0026#39;s API serves the model at full precision for setups that cannot run it locally.\u003c/p\u003e","blueskyPost":"FlashLabs shipped GLM-5.3-MLX, an MLX quantization of Z.ai's 743B-parameter GLM-5.3 for Apple Silicon. OrcaSAQ cuts the 2-bit build from ~1.51TB to ~322GB while keeping ~80.2% accuracy. A 512GB Mac runs the 2-4 bit variants locally. GGUF via Unsloth.","twitterPost":"FlashLabs released GLM-5.3-MLX, an MLX build of Z.ai's 743B GLM-5.3 for Apple Silicon. The 2-bit variant shrinks ~1.51TB to ~322GB (~79%) at ~80.2% accuracy. 512GB Macs run 2-4 bit locally; 6-bit needs two Macs or 8x H200. GGUF from Unsloth.","threadsPost":null,"newsletterBlurb":"FlashLabs published GLM-5.3-MLX, an MLX quantization of Z.ai's 743B-parameter GLM-5.3, on Hugging Face. The OrcaSAQ method cuts the 2-bit build from about 1.51TB to roughly 322GB while keeping about 80.2 percent accuracy, and a 512GB Mac with M3 Ultra runs the 2-bit through 4-bit variants locally. Unsloth also provides a GGUF version for llama.cpp environments.","attributionJson":"[{\"source\":\"ASCII.jp\",\"url\":\"https://ascii.jp/elem/000/004/430/4430810/?rss\",\"title\":\"OrcaRouter、1.5TB超の大規模LLM「GLM-5.3」を1台のMacで実行できる「GLM-5.3-MLX」を公開\"}]","lintFlagsJson":null,"lintHits":0,"costUsd":0,"inputTokens":5712,"outputTokens":2992,"status":"published","repairAttempts":0,"nextRepairAt":null,"factsAttemptedAt":1788170127,"createdAt":"2026-08-31T09:47:25.000Z","publishedAt":"2026-08-31T09:47:28.000Z","updatedAt":"2026-08-31T09:47:28.000Z"},"cluster":{"id":"c_e68ef8a1d06137189d00845f","canonicalTitle":"OrcaRouter、1.5TB超の大規模LLM「GLM-5.3」を1台のMacで実行できる「GLM-5.3-MLX」を公開","representativeArticleId":"a_050b5e5ed233cc20493a05f7","sourceCount":1,"writtenSourceCount":1,"writeAttempts":0,"isSolo":true,"entitiesJson":"{\"anime_titles\":[],\"manga_titles\":[],\"work_titles\":[],\"studios\":[],\"people\":[],\"type\":\"announcement\",\"domain\":\"other\",\"is_roundup\":false}","contentType":"news","status":"published","firstSeenAt":"2026-08-31T05:30:02.000Z","lastSeenAt":"2026-08-31T05:30:02.000Z","updatedAt":"2026-08-31T09:47:28.000Z"},"attribution":[{"source":"ASCII.jp","url":"https://ascii.jp/elem/000/004/430/4430810/?rss","title":"OrcaRouter、1.5TB超の大規模LLM「GLM-5.3」を1台のMacで実行できる「GLM-5.3-MLX」を公開"}],"entities":{"anime_titles":[],"manga_titles":[],"work_titles":[],"studios":[],"people":[],"type":"announcement","domain":"other","is_roundup":false},"keyFacts":null}
