{"rewrite":{"id":"r_707d3669059b9201f4c1436d","clusterId":"c_81326f06c7ffb2935cdee682","slug":"nishika-open-sources-j-meeteval-for-japanese-meeting-summaries","model":"deepseek-v4-1-flash","headline":"Nishika Open-Sources J-MeetEval For Japanese Meeting Summaries","summary":"Nishika released J-MeetEval, a benchmark for instruction following in Japanese meeting summarization, as open source on GitHub and Hugging Face. The dataset holds 197 synthetic meeting transcripts, 52 instruction definitions, 408 total instructions and 27 instruction types across 16 categories, with inputs of 3,000 to 11,000 characters. Scoring is binary per instruction, judged by gpt-5-mini with a three-vote majority. Across seven models in the 4B to 9B class, scores on the Nejumi Leaderboard ran opposite to pass rates here (Spearman rho of minus 0.52).","whyItMatters":"General-purpose leaderboard rank does not predict whether a Japanese model will follow formatting and scope instructions on long meeting transcripts, which is the gap Nishika built a task-specific benchmark to measure.","webCardHtml":"\u003cp\u003eJ-MeetEval sets several instructions at once on a single transcript, such as unifying numerals to full-width characters, sorting by speaker order, and returning a fixed JSON shape, then scores each one pass or fail. Formatting results split by direction: rounding numerals to half-width passed 98 percent of the time, while the opposite demand for full-width managed 27 percent. Instructions that narrow scope were harder still, with extracting only decisions and their supporting remarks at 25 percent. Combining instructions also hurt, with chronological ordering dropping sharply when paired with anonymization, job titles or JSON output.\u003c/p\u003e","blueskyPost":"J-MeetEval's scoring pairs a binary judge with three-vote majority, so a model that writes a polished minuted misses a single buried instruction and nets zero on it. That is why its pass rates invert the Nejumi ranking.","twitterPost":"J-MeetEval scores each instruction binary via a three-vote majority, so one missed buried instruction zeroes it. Hence the inverted ranking.","threadsPost":"J-MeetEval grades every instruction on a binary, judged by gpt-5-mini with a three-vote majority. A fluent summary loses the same credit for a buried instruction as it earns for all the easy ones, which is why the 4B to 9B models' pass rates run opposite to their Nejumi Leaderboard standing.","newsletterBlurb":"Nishika released J-MeetEval, an open-source benchmark that tests whether Japanese models follow formatting and scope instructions on long meeting transcripts rather than whether they write well. It holds 197 synthetic transcripts, 52 instruction definitions and 408 total instructions, scored pass or fail. Across seven models in the 4B to 9B class, Nejumi Leaderboard rank and J-MeetEval pass rates ran opposite, a correlation of minus 0.52.","attributionJson":"[{\"source\":\"ASCII.jp\",\"url\":\"https://ascii.jp/elem/000/004/440/4440121/?rss\",\"title\":\"【GENIAC第3期成果】日本語の会議要約AI向け評価ベンチマーク「J-MeetEval」をオープンソースで公開。汎用ベンチマークの順位と\\\"議事録タスクでの実力\\\"が一致しないことを実証。\"}]","lintFlagsJson":null,"lintHits":0,"costUsd":0,"inputTokens":6065,"outputTokens":698,"status":"published","repairAttempts":0,"nextRepairAt":null,"factsAttemptedAt":1791197962,"createdAt":"2026-10-05T10:54:08.000Z","publishedAt":"2026-10-05T10:58:20.000Z","updatedAt":"2026-10-05T10:58:20.000Z"},"cluster":{"id":"c_81326f06c7ffb2935cdee682","canonicalTitle":"【GENIAC第3期成果】日本語の会議要約AI向け評価ベンチマーク「J-MeetEval」をオープンソースで公開。汎用ベンチマークの順位と\"議事録タスクでの実力\"が一致しないことを実証。","representativeArticleId":"a_880a9a86e4468db14786b4aa","sourceCount":1,"writtenSourceCount":1,"writeAttempts":0,"isSolo":true,"entitiesJson":"{\"anime_titles\":[],\"manga_titles\":[],\"work_titles\":[\"J-MeetEval\"],\"studios\":[],\"people\":[],\"type\":\"announcement\",\"domain\":\"other\",\"is_roundup\":false}","contentType":"news","status":"published","firstSeenAt":"2026-10-05T06:29:00.000Z","lastSeenAt":"2026-10-05T06:29:00.000Z","updatedAt":"2026-10-05T10:58:21.000Z"},"attribution":[{"source":"ASCII.jp","url":"https://ascii.jp/elem/000/004/440/4440121/?rss","title":"【GENIAC第3期成果】日本語の会議要約AI向け評価ベンチマーク「J-MeetEval」をオープンソースで公開。汎用ベンチマークの順位と\"議事録タスクでの実力\"が一致しないことを実証。"}],"entities":{"anime_titles":[],"manga_titles":[],"work_titles":["J-MeetEval"],"studios":[],"people":[],"type":"announcement","domain":"other","is_roundup":false},"keyFacts":null}
