{"rewrite":{"id":"r_c755ef20eb579166c18efbe8","clusterId":"c_d29029d2af06c396753d8d54","slug":"openai-releases-lifescibench-benchmark-for-life-science-ai","model":"deepseek-v4-flash","headline":"OpenAI Releases LifeSciBench Benchmark for Life Science AI","summary":"OpenAI announced LifeSciBench, a benchmark test measuring how useful AI is for life science researchers. Developed with 173 scientists, it includes 750 tasks and 1,062 attachments, evaluating AI on criteria like reasoning and detail level. OpenAI also released a report on GPT-5.4 assisting drug discovery.","whyItMatters":"LifeSciBench shifts AI evaluation from narrow question-and-answer tests to real-world scientific workflows, setting a new standard for measuring practical utility in research.","webCardHtml":"\u003cp\u003eOpenAI released LifeSciBench on June 17, a benchmark designed to measure how useful AI is for life science researchers. Unlike traditional science tests that focus on narrow domain knowledge or clear-answer Q\u0026amp;A formats, LifeSciBench classifies daily scientist tasks into seven categories including handling scientific evidence, analysis, design and optimization, and scientific communication. OpenAI worked with 173 scientists in biotechnology and drug discovery to create the tasks, each structured as a scientist requesting help from a knowledgeable collaborator.\u003c/p\u003e\u003cp\u003eThe benchmark gives AI models 750 tasks with 1,062 attachments such as figures, tables, and chemical structure files, with 53% of tasks requiring at least one attachment. Responses are scored on criteria including appropriate detail, correct reasoning, and correct format. The highest-scoring model was GPT-Rosalind, OpenAI\u0026#39;s science-specialized AI based on GPT-5.5, which outperformed GPT-5.5 across all seven categories. On the same day, OpenAI also published a report on GPT-5.4 assisting drug discovery research.\u003c/p\u003e","blueskyPost":"OpenAI released LifeSciBench, a benchmark measuring how useful AI is for life scientists. 750 tasks, 1062 attachments, scored on real-world criteria. GPT-Rosalind topped the rankings.","twitterPost":"OpenAI released LifeSciBench, a benchmark for life science AI usefulness. 750 tasks, 1062 attachments, scored on real-world criteria. GPT-Rosalind topped the rankings.","threadsPost":null,"newsletterBlurb":"OpenAI announced LifeSciBench, a benchmark test that measures how useful AI is for life science researchers, developed with 173 scientists. It includes 750 tasks with attachments and evaluates AI on criteria like reasoning and detail. GPT-Rosalind scored highest, and OpenAI also reported GPT-5.4 assisted drug discovery.","attributionJson":"[{\"source\":\"GIGAZINE\",\"url\":\"https://gigazine.net/news/20260618-openai-lifescibench/\",\"title\":\"OpenAI Releases 'LifeSciBench,' a Benchmark Test That Measures How Useful AI Is for Scientists\"}]","lintFlagsJson":null,"lintHits":0,"costUsd":0,"inputTokens":4206,"outputTokens":595,"status":"published","repairAttempts":0,"nextRepairAt":null,"factsAttemptedAt":1781751960,"createdAt":"2026-06-18T03:01:25.000Z","publishedAt":"2026-06-18T03:04:57.000Z","updatedAt":"2026-06-18T03:04:57.000Z"},"cluster":{"id":"c_d29029d2af06c396753d8d54","canonicalTitle":"AIが科学者にとってどれだけ役立つかを測定できるベンチマークテスト「LifeSciBench」をOpenAIが公開","representativeArticleId":"a_206cbdc7eddadb139dcf9e64","sourceCount":1,"writtenSourceCount":1,"writeAttempts":0,"isSolo":true,"entitiesJson":"{\"anime_titles\":[],\"manga_titles\":[],\"work_titles\":[],\"studios\":[],\"people\":[],\"type\":\"news\",\"domain\":\"other\",\"is_roundup\":false}","contentType":"news","status":"published","firstSeenAt":"2026-06-18T02:17:00.000Z","lastSeenAt":"2026-06-18T02:17:00.000Z","updatedAt":"2026-06-18T03:04:58.000Z"},"attribution":[{"source":"GIGAZINE","url":"https://gigazine.net/news/20260618-openai-lifescibench/","title":"AIが科学者にとってどれだけ役立つかを測定できるベンチマークテスト「LifeSciBench」をOpenAIが公開"}],"entities":{"anime_titles":[],"manga_titles":[],"work_titles":[],"studios":[],"people":[],"type":"news","domain":"other","is_roundup":false},"keyFacts":["OpenAI released LifeSciBench on June 17, a benchmark with 750 tasks and 1,062 attachments designed to measure AI utility for life science researchers.","The highest-scoring model on LifeSciBench was GPT-Rosalind, OpenAI's science-specialized AI based on GPT-5.5, which outperformed GPT-5.5 across all seven categories.","OpenAI also published a report on GPT-5.4 assisting drug discovery research on the same day as the LifeSciBench release."]}
