{"rewrite":{"id":"r_924f6f3e0a085fb274ec7845","clusterId":"c_12ad3eca75aacf22afea49ae","slug":"hayamimi-runs-real-time-multilingual-speech-recognition-on-cpu-alone","model":"deepseek-v4-1-flash","headline":"Hayamimi Runs Real-Time Multilingual Speech Recognition on CPU Alone","summary":"Hayamimi is a real-time multilingual speech recognition system that runs on CPU without GPUs or cloud APIs. It starts showing subtitles while speech is in progress and finalizes Japanese output about 100ms after a speaker stops. It works under 2GB of memory, supports live subtitles, speaker labeling, and translated subtitles, and uses INT8-quantized ONNX models on sherpa-onnx.","whyItMatters":"A CPU-only pipeline that hits 3.8% character error rate on Japanese broadcast audio removes the GPU and cloud dependency that most real-time speech recognition still assumes, which puts live subtitles within reach on ordinary hardware.","webCardHtml":"\u003cp\u003eHayamimi routes each utterance to a dedicated model after determining its language, an approach that departs from the single-model setup typical of CPU-based speech recognition. The models are INT8-quantized ONNX files running on sherpa-onnx, so PyTorch and CUDA are not required. On a 6-core desktop CPU it reaches speeds 10 to 50 times real time. A two-pass correction step re-decodes the preceding utterance after two seconds of silence, which improves Japanese character error rate from 15.5% to 12.0%. It requires Python 3.10 or higher and ffmpeg on the PATH.\u003c/p\u003e","blueskyPost":"Hayamimi does real-time multilingual speech recognition on CPU alone, no GPU or cloud API. Under 2GB of memory. Japanese subtitles finalize about 100ms after a speaker stops, with 3.8% character error rate on broadcast audio.","twitterPost":"Hayamimi runs real-time multilingual speech recognition on CPU only, no GPU or cloud. Under 2GB memory, subtitles finalize about 100ms after speech ends, 3.8% CER on Japanese broadcast audio via INT8 ONNX on sherpa-onnx.","threadsPost":null,"newsletterBlurb":"Hayamimi is a real-time multilingual speech recognition system that runs on CPU without GPUs or cloud APIs, holding under 2GB of memory. It shows draft subtitles during speech and finalizes Japanese output about 100ms after a speaker stops, reaching a 3.8% character error rate on Japanese TV broadcast audio. Speaker labeling, translation, and a browser dashboard are included.","attributionJson":"[{\"source\":\"GIGAZINE\",\"url\":\"https://gigazine.net/news/20261011-hayamimi/\",\"title\":\"CPU-only real-time multilingual speech recognition 'Hayamimi' works without GPU or cloud APIs, runs under 2GB of memory, and supports everything from live subtitles to browser display, speaker labeling, and translated subtitles\"}]","lintFlagsJson":null,"lintHits":0,"costUsd":0,"inputTokens":4696,"outputTokens":619,"status":"published","repairAttempts":0,"nextRepairAt":null,"factsAttemptedAt":1791692536,"createdAt":"2026-10-11T04:17:27.000Z","publishedAt":"2026-10-11T04:21:08.000Z","updatedAt":"2026-10-11T04:21:08.000Z"},"cluster":{"id":"c_12ad3eca75aacf22afea49ae","canonicalTitle":"CPUのみで動作するリアルタイム多言語音声認識「早耳」、GPUもクラウドAPIも使わずメモリ2GB未満でライブ字幕表示からブラウザ表示・話者ラベル付け・翻訳字幕まで可能","representativeArticleId":"a_a5273d1003061e6665156f63","sourceCount":1,"writtenSourceCount":1,"writeAttempts":0,"isSolo":true,"entitiesJson":"{\"anime_titles\":[],\"manga_titles\":[],\"work_titles\":[\"Hayamimi\"],\"studios\":[],\"people\":[],\"type\":\"news\",\"domain\":\"other\",\"is_roundup\":false}","contentType":"news","status":"published","firstSeenAt":"2026-10-11T03:00:00.000Z","lastSeenAt":"2026-10-11T03:00:00.000Z","updatedAt":"2026-10-11T04:21:08.000Z"},"attribution":[{"source":"GIGAZINE","url":"https://gigazine.net/news/20261011-hayamimi/","title":"CPUのみで動作するリアルタイム多言語音声認識「早耳」、GPUもクラウドAPIも使わずメモリ2GB未満でライブ字幕表示からブラウザ表示・話者ラベル付け・翻訳字幕まで可能"}],"entities":{"anime_titles":[],"manga_titles":[],"work_titles":["Hayamimi"],"studios":[],"people":[],"type":"news","domain":"other","is_roundup":false},"keyFacts":null}
