{"rewrite":{"id":"r_59e13cfa0d0b092a7fe9b968","clusterId":"c_8c9e51a1e629928ee73fe282","slug":"ltx-2-5-lip-sync-pipeline-cuts-reply-wait-to-2-46-seconds-per-chunk","model":"deepseek-v4-1-flash","headline":"LTX-2.5 Lip-Sync Pipeline Cuts Reply Wait To 2.46 Seconds Per Chunk","summary":"A writer developing a real-time talking avatar of his late wife reported that a GitHub update to the LTX-2.5 accelerated video generation stack cut single-chunk generation time from 4.45 seconds to 2.46 seconds. The system splits spoken replies into chunks of 4.8 seconds or less and generates lip-synced video for each. Earlier configurations took 14.6 seconds or more before a reply appeared.","whyItMatters":"The reporter's own benchmarks show the gap between generated video and speech playback was the wall keeping avatar conversation from feeling like conversation, and layer streaming in Transformer Engine is what moved a 32GB consumer GPU past it.","webCardHtml":"\u003cp\u003eThe pipeline splits a reply\u0026#39;s audio into chunks of 4.8 seconds or shorter, cutting at the quietest point between 3.6 and 4.8 seconds, and returns lip-synced video for each chunk. Playback starts as soon as the first chunk arrives while the rest generate. A chaining step that makes the first frame of each later chunk match the last frame of the previous one reduces pose jumps at the seams by a third.\u003c/p\u003e\u003cp\u003eOn a 32GB GPU, one chunk took 4.45 seconds before the update and 2.46 seconds after. An earlier two-model setup using MiniMax H3 with TaoMate had the highest lip-sync accuracy of the configurations tried, but the wait from speaking to a reply was about 35 seconds.\u003c/p\u003e","blueskyPost":"LTX-2.5 chunked lip-sync generation dropped from 4.45 to 2.46 seconds per chunk on a 32GB GPU, per a developer's own benchmarks. The prior full-reply approach took 14.6 seconds to surface a response.","twitterPost":"Chunked LTX-2.5 lip-sync generation cut per-chunk time from 4.45 to 2.46 seconds on a 32GB GPU. The earlier whole-reply setup took 14.6 seconds before a reply appeared.","threadsPost":null,"newsletterBlurb":"A developer building a real-time talking avatar reported that a layer-streaming update to the LTX-2.5 stack cut single-chunk lip-sync generation from 4.45 seconds to 2.46 seconds on a 32GB GPU. The system splits replies into chunks of 4.8 seconds or less so playback can start before the full response finishes generating. The previous whole-reply approach took 14.6 seconds before a reply appeared.","attributionJson":"[{\"source\":\"GameBusiness.jp\",\"url\":\"https://www.gamebusiness.jp/article/2026/09/29/28156.html\",\"title\":\"話しかけて6秒で、彼女が口を動かして返事する。LTX-2.5高速化技術で、動画生成リップシンクアバターがようやく「会話」になった（CloseBox）\"}]","lintFlagsJson":null,"lintHits":0,"costUsd":0,"inputTokens":6054,"outputTokens":689,"status":"published","repairAttempts":0,"nextRepairAt":null,"factsAttemptedAt":1791234425,"createdAt":"2026-10-05T21:03:02.000Z","publishedAt":"2026-10-05T21:06:05.000Z","updatedAt":"2026-10-05T21:06:05.000Z"},"cluster":{"id":"c_8c9e51a1e629928ee73fe282","canonicalTitle":"話しかけて6秒で、彼女が口を動かして返事する。LTX-2.5高速化技術で、動画生成リップシンクアバターがようやく「会話」になった（CloseBox）","representativeArticleId":"a_2a1eb29e9447fe7945c177d7","sourceCount":1,"writtenSourceCount":1,"writeAttempts":0,"isSolo":true,"entitiesJson":"{\"anime_titles\":[],\"manga_titles\":[],\"work_titles\":[],\"studios\":[],\"people\":[],\"type\":\"news\",\"domain\":\"other\",\"is_roundup\":false}","contentType":"news","status":"published","firstSeenAt":"2026-09-28T21:15:03.000Z","lastSeenAt":"2026-09-28T21:15:03.000Z","updatedAt":"2026-10-05T21:06:06.000Z"},"attribution":[{"source":"GameBusiness.jp","url":"https://www.gamebusiness.jp/article/2026/09/29/28156.html","title":"話しかけて6秒で、彼女が口を動かして返事する。LTX-2.5高速化技術で、動画生成リップシンクアバターがようやく「会話」になった（CloseBox）"}],"entities":{"anime_titles":[],"manga_titles":[],"work_titles":[],"studios":[],"people":[],"type":"news","domain":"other","is_roundup":false},"keyFacts":null}
