{"rewrite":{"id":"r_c95189a269d83ed43e1a6b11","clusterId":"c_55d8adaebc62b0755d114e50","slug":"ai-music-video-took-34-hours-and-many-rejected-takes","model":"deepseek-v4-1-flash","headline":"Ai Music Video Took 34 Hours And Many Rejected Takes","summary":"A column from ASCII.jp describes an attempt to build a full music video with GPT-6 Astra, released September 3, and the local video model MiniMax H3. The song, 'Design and Will,' came from Suno v6, which arrived September 9. The author reports 34 hours of production, repeated rejections of generated output, and a conclusion that a human still has to direct the AI closely to keep it on course.","whyItMatters":"The column's own account suggests AI video tools now handle complex staging only when a person steers them, which puts the near-term pressure on directing skill rather than on the models alone.","webCardHtml":"\u003cp\u003eThe author of the ASCII.jp column says the song came first. \u0026#39;Design and Will\u0026#39; was built with Suno v6, the music model released on September 9, and the lyrics were produced by having GPT-6 Astra summarize a conversation into verse. From there the column describes trying to extend that into a full music video, using MiniMax H3 for video generation in a local environment.\u003c/p\u003e\u003cp\u003eThe stated result is 34 hours of work and repeated rejections, with output that sometimes drifted from the intended direction until the author stepped in to direct it. The column\u0026#39;s own reading is that human control is what makes the complex staging possible, and that AI may push media further toward personalization.\u003c/p\u003e","blueskyPost":"The 34 hours went into rejecting output, not rendering it. When the bottleneck is judgment rather than compute, AI video tools shift the scarce skill from animating to directing.","twitterPost":"34 hours of production, most of it spent rejecting takes. The scarce skill stops being animation and becomes direction.","threadsPost":"Most of the 34 hours went to rejecting generated takes, not rendering them. That reframes what AI video tools change: the scarce skill shifts from animating to directing, since a human still has to catch what the model got wrong.","newsletterBlurb":"An ASCII.jp column recounts building a music video for 'Design and Will' with GPT-6 Astra, released September 3, and the local video model MiniMax H3. The author reports 34 hours of production and repeated rejections of generated footage. The stated conclusion is that close human direction is what makes complex staging possible, and that AI may push media further toward personalization.","attributionJson":"[{\"source\":\"ASCII.jp\",\"url\":\"https://ascii.jp/elem/000/004/436/4436544/?rss\",\"title\":\"Making an MV with AI Was Harder Than I Imagined: 34 Hours of Production and Repeated 'Rejections' Taught Me Something\"}]","lintFlagsJson":null,"lintHits":0,"costUsd":0,"inputTokens":4702,"outputTokens":597,"status":"published","repairAttempts":0,"nextRepairAt":null,"factsAttemptedAt":1789946960,"createdAt":"2026-09-20T23:23:41.000Z","publishedAt":"2026-09-20T23:28:20.000Z","updatedAt":"2026-09-20T23:28:20.000Z"},"cluster":{"id":"c_55d8adaebc62b0755d114e50","canonicalTitle":"AIでMVを作るのは想像以上に大変だった　制作34時間、繰り返した“ダメ出し”でわかったこと","representativeArticleId":"a_79c297cd1690a6702368ffbb","sourceCount":1,"writtenSourceCount":1,"writeAttempts":0,"isSolo":true,"entitiesJson":"{\"anime_titles\":[],\"manga_titles\":[],\"work_titles\":[],\"studios\":[],\"people\":[],\"type\":\"other\",\"domain\":\"music\",\"is_roundup\":false}","contentType":"news","status":"published","firstSeenAt":"2026-09-20T22:00:00.000Z","lastSeenAt":"2026-09-20T22:00:00.000Z","updatedAt":"2026-09-20T23:28:20.000Z"},"attribution":[{"source":"ASCII.jp","url":"https://ascii.jp/elem/000/004/436/4436544/?rss","title":"AIでMVを作るのは想像以上に大変だった　制作34時間、繰り返した“ダメ出し”でわかったこと"}],"entities":{"anime_titles":[],"manga_titles":[],"work_titles":[],"studios":[],"people":[],"type":"other","domain":"music","is_roundup":false},"keyFacts":null}
