{"rewrite":{"id":"r_07d5a5b2b30dbf9ae36bb83e","clusterId":"c_f950b273c0352c64b3cce6fd","slug":"opentpu-project-runs-language-models-on-kintex-7-fpga-card","model":"deepseek-v4-1-flash","headline":"OpenTPU Project Runs Language Models On Kintex-7 FPGA Card","summary":"An open-source project called openTPU has been released, built around a simple accelerator design that pairs a sequencer, DMA, matrix and vector units, and a quantization block. The developer wrote the circuits onto an FPGA card using AMD's Kintex-7 and ran LFM2.5-230M, Qwen3-0.6B, Qwen3.5-0.8B, Gemma 4, and Phi-4-mini, including the 34.7-billion-parameter Qwen3.5-35B-A3B. It ships under the Apache License 2.0.","whyItMatters":"The stated aim is not to reproduce Google's TPU but to test how far AI agents can carry hardware design, including the design of the chips that run them, and the whole stack is published for inspection.","webCardHtml":"\u003cp\u003eThe repository bundles the RTL, instruction set architecture, simulator, compiler, and profiler in one place, which is what makes the design inspectable rather than a black box. Its structure is deliberately plain: a sequencer issuing instructions in order, a DMA unit moving data to and from memory, a matrix operation unit, a vector unit for floating-point work, and a quantization stage. There are no caches and no complex instruction scheduling of the kind general CPUs and GPUs carry, and data movement is written out as instructions, so it stays possible to trace which process consumed how many clocks. The developer cites memory read speed for model data, not the arithmetic itself, as the current constraint.\u003c/p\u003e","blueskyPost":"openTPU is out: an open-source AI accelerator whose design was handled by AI agents. Runs LFM2.5-230M, Qwen3-0.6B, Qwen3.5-0.8B, Gemma 4 and Phi-4-mini on a Kintex-7 FPGA card, plus Qwen3.5-35B-A3B at 3.95 tokens/sec. Apache 2.0.","twitterPost":"openTPU released: an open-source AI accelerator where AI handled the design process. LFM2.5-230M at 82.1 tokens/sec, Qwen3-0.6B at 30.7, Qwen3.5-0.8B at 23.3 on a Kintex-7 FPGA card. Qwen3.5-35B-A3B also ran, at 3.95 tokens/sec.","threadsPost":null,"newsletterBlurb":"openTPU, an open-source AI accelerator developed with AI agents doing the design work, has been released. The developer ran it on an FPGA card with AMD's Kintex-7, covering LFM2.5-230M, Qwen3-0.6B, Qwen3.5-0.8B, Gemma 4, and Phi-4-mini, and also reports running the 34.7-billion-parameter Qwen3.5-35B-A3B by streaming only needed data from the PC. Everything from software to circuits is published under the Apache License 2.0.","attributionJson":"[{\"source\":\"GIGAZINE\",\"url\":\"https://gigazine.net/news/20261007-opentpu/\",\"title\":\"AI-developed open-source AI accelerator \\\"openTPU\\\" emerges\"}]","lintFlagsJson":null,"lintHits":0,"costUsd":0,"inputTokens":4749,"outputTokens":762,"status":"published","repairAttempts":0,"nextRepairAt":null,"factsAttemptedAt":1791370338,"createdAt":"2026-10-07T10:47:41.000Z","publishedAt":"2026-10-07T10:51:10.000Z","updatedAt":"2026-10-07T10:51:10.000Z"},"cluster":{"id":"c_f950b273c0352c64b3cce6fd","canonicalTitle":"AIによって開発されたオープンソースのAIアクセラレータ「openTPU」が登場","representativeArticleId":"a_38d9ef916f3b04f528a88ffc","sourceCount":1,"writtenSourceCount":1,"writeAttempts":0,"isSolo":true,"entitiesJson":"{\"anime_titles\":[],\"manga_titles\":[],\"work_titles\":[],\"studios\":[],\"people\":[],\"type\":\"news\",\"domain\":\"other\",\"is_roundup\":false}","contentType":"news","status":"published","firstSeenAt":"2026-10-07T10:00:00.000Z","lastSeenAt":"2026-10-07T10:00:00.000Z","updatedAt":"2026-10-07T10:51:07.000Z"},"attribution":[{"source":"GIGAZINE","url":"https://gigazine.net/news/20261007-opentpu/","title":"AIによって開発されたオープンソースのAIアクセラレータ「openTPU」が登場"}],"entities":{"anime_titles":[],"manga_titles":[],"work_titles":[],"studios":[],"people":[],"type":"news","domain":"other","is_roundup":false},"keyFacts":null}
