{"rewrite":{"id":"r_0ea616c37e18ce20b4497f39","clusterId":"c_9dedc68b353498fcfbf07bdd","slug":"openai-publishes-research-on-training-ai-in-honesty-and-humility","model":"deepseek-v4-flash","headline":"OpenAI Publishes Research on Training Ai in Honesty and Humility","summary":"OpenAI published research on June 18 showing that training AI in beneficial traits like honesty, admitting uncertainty, and accepting correction leads to those behaviors spreading to untrained areas and improving resistance to malicious instructions. The study used reinforcement learning with 15 traits across 12 fields and found the trained AI outperformed a standard model in 44 of 53 evaluations.","whyItMatters":"The research suggests that instilling traits like honesty and humility in AI through reinforcement learning can produce broadly beneficial behavior that generalizes beyond training data and resists manipulation.","webCardHtml":"\u003cp\u003eOpenAI published research on June 18 showing that training AI in beneficial traits such as honesty, humility in admitting uncertainty, openness to correction, and fairness leads to desirable behavior spreading to untrained areas and becoming more resistant to malicious instructions. The study, titled \u0026#39;Reinforcement learning towards broadly and persistently beneficial models,\u0026#39; used reinforcement learning with 15 traits across 12 fields including healthcare, education, science, law, engineering, and economics.\u003c/p\u003e\u003cp\u003eThe research team trained AI using 95% standard reinforcement learning data and 5% data for learning beneficial traits, then compared it with AI trained on standard data alone. The AI that learned beneficial traits outperformed the comparison in 44 out of 53 evaluations prepared separately. The study also found that behavior changed beyond the learned fields: AI trained only with additional healthcare conversations showed improvement in 17 evaluations unrelated to healthcare, such as reward hacking and deception in programming.\u003c/p\u003e","blueskyPost":"OpenAI published research showing that training AI in honesty and humility through reinforcement learning makes those behaviors spread to untrained areas and resist malicious instructions. The trained AI outperformed a standard model in 44 of 53 evaluations.","twitterPost":"OpenAI published research showing that training AI in honesty and humility through reinforcement learning makes those behaviors spread to untrained areas and resist malicious instructions. The trained AI outperformed a standard model in 44 of 53 evaluations.","threadsPost":null,"newsletterBlurb":"OpenAI published research on June 18 showing that training AI in beneficial traits like honesty, admitting uncertainty, and accepting correction leads to those behaviors spreading to untrained areas and improving resistance to malicious instructions. The study used reinforcement learning with 15 traits across 12 fields and found the trained AI outperformed a standard model in 44 of 53 evaluations.","attributionJson":"[{\"source\":\"GIGAZINE\",\"url\":\"https://gigazine.net/news/20260619-openai-beneficial-rl/\",\"title\":\"Can AI acquire the ability to admit when it doesn't know something? OpenAI publishes research results on instilling beneficial traits through reinforcement learning\"}]","lintFlagsJson":null,"lintHits":0,"costUsd":0,"inputTokens":4074,"outputTokens":601,"status":"published","repairAttempts":0,"nextRepairAt":null,"factsAttemptedAt":1781871683,"createdAt":"2026-06-19T12:14:32.000Z","publishedAt":"2026-06-19T12:18:15.000Z","updatedAt":"2026-06-19T12:18:15.000Z"},"cluster":{"id":"c_9dedc68b353498fcfbf07bdd","canonicalTitle":"AIに「分からないことを分からないと認める力」は身につくのか？OpenAIが有益な性質を強化学習で定着させる研究結果を公開","representativeArticleId":"a_97419879e3e1dd6f0bba5295","sourceCount":1,"writtenSourceCount":1,"writeAttempts":0,"isSolo":true,"entitiesJson":"{\"anime_titles\":[],\"manga_titles\":[],\"work_titles\":[],\"studios\":[],\"people\":[],\"type\":\"news\",\"domain\":\"other\",\"is_roundup\":false}","contentType":"news","status":"published","firstSeenAt":"2026-06-19T11:00:00.000Z","lastSeenAt":"2026-06-19T11:00:00.000Z","updatedAt":"2026-06-19T12:18:16.000Z"},"attribution":[{"source":"GIGAZINE","url":"https://gigazine.net/news/20260619-openai-beneficial-rl/","title":"AIに「分からないことを分からないと認める力」は身につくのか？OpenAIが有益な性質を強化学習で定着させる研究結果を公開"}],"entities":{"anime_titles":[],"manga_titles":[],"work_titles":[],"studios":[],"people":[],"type":"news","domain":"other","is_roundup":false},"keyFacts":["OpenAI published research on June 18 showing that training AI in beneficial traits like honesty, admitting uncertainty, and accepting correction leads to those behaviors spreading to untrained areas.","The study used reinforcement learning with 15 traits across 12 fields including healthcare, education, science, law, engineering, and economics.","The AI trained with beneficial traits outperformed a standard model in 44 of 53 evaluations.","AI trained only with additional healthcare conversations showed improvement in 17 evaluations unrelated to healthcare, such as reward hacking and deception in programming."]}
