{"rewrite":{"id":"r_4755819a9ef71f63447f0d8a","clusterId":"c_7217a20ca77cc0b7655e4359","slug":"nvidia-agent-avo-scores-100-on-arc-agi-3-where-base-model-manages-30","model":"deepseek-v4-flash","headline":"NVIDIA Agent AVO Scores 100% on ARC-AGI-3 Where Base Model Manages 30%","summary":"NVIDIA's Agentic Variation Operators (AVO) agent system scored 100% on the public set of ARC-AGI-3, a benchmark measuring AI reasoning in unknown game-like environments. The base model 5 used inside AVO scored about 30% in a separate evaluation by ARC Prize, the benchmark's developer. AVO cleared all 183 levels across 25 environment types.","whyItMatters":"The gap between the base model's standalone 30% and the full agent's 100% shows that the execution harness handling memory, tools, and recovery determines long-horizon performance more than the AI model itself.","webCardHtml":"\u003cp\u003eNVIDIA reports that AVO, its agent system built originally for GPU kernel optimization, cleared all 183 levels in the ARC-AGI-3 public set while keeping action efficiency equal to or better than first-time human players. The benchmark\u0026#39;s metric, Relative Human Action Efficiency, measures both clearing the game and how few actions the agent takes compared to a human baseline.\u003c/p\u003e\u003cp\u003eARC Prize, which developed ARC-AGI-3, evaluated the same base model under different conditions and recorded roughly 30% on the same metric. The same agent architecture transferred from GPU code optimization to unknown games without redesign. AVO uses persistent memory of past trials and a supervision function that redirects strategy when progress stalls.\u003c/p\u003e","blueskyPost":"NVIDIA AVO cleared all 183 ARC-AGI-3 levels while its base model 5 scored 30% in ARC Prize's separate run. The harness, not the model, carries the reasoning gain.","twitterPost":"NVIDIA AVO hit 100% on ARC-AGI-3 while base model 5 scored 30% separately. The harness is the performance.","threadsPost":"NVIDIA AVO scored 100% on ARC-AGI-3's public set, clearing all 183 levels. Its base model 5 managed 30% in ARC Prize's separate evaluation. The gap isolates the harness as the decisive component.","newsletterBlurb":"NVIDIA's AVO agent system recorded 100% on the ARC-AGI-3 public set, clearing all 183 levels. ARC Prize's separate evaluation of the same base model scored about 30%, pointing to the execution harness, not the model, as the deciding factor in long-horizon tasks.","attributionJson":"[{\"source\":\"GIGAZINE\",\"url\":\"https://gigazine.net/news/20260825-nvidia-avo/\",\"title\":\"NVIDIA's AI Agent 'AVO' Achieves 100% on ARC-AGI-3, While Evaluation of the Same Base Model Under Different Conditions Is About 30%, Highlighting the Importance of the 'Execution Harness'\"}]","lintFlagsJson":null,"lintHits":0,"costUsd":0,"inputTokens":4664,"outputTokens":628,"status":"published","repairAttempts":0,"nextRepairAt":null,"factsAttemptedAt":1788002277,"createdAt":"2026-08-29T11:15:56.000Z","publishedAt":"2026-08-29T11:16:44.000Z","updatedAt":"2026-08-29T11:15:56.000Z"},"cluster":{"id":"c_7217a20ca77cc0b7655e4359","canonicalTitle":"NVIDIAのAIエージェント「AVO」がARC-AGI-3で100％を達成、同じベースモデルの別条件での評価は約30％で「実行基盤(ハーネス)」の重要性が示される","representativeArticleId":"a_ef90166c03e2f9aad5311d91","sourceCount":1,"writtenSourceCount":1,"writeAttempts":1,"isSolo":true,"entitiesJson":"{\"anime_titles\":[],\"manga_titles\":[],\"work_titles\":[\"ARC-AGI-3\"],\"studios\":[\"NVIDIA\"],\"people\":[],\"type\":\"news\",\"domain\":\"other\",\"is_roundup\":false}","contentType":"news","status":"published","firstSeenAt":"2026-08-24T23:00:00.000Z","lastSeenAt":"2026-08-24T23:00:00.000Z","updatedAt":"2026-08-29T11:16:46.000Z"},"attribution":[{"source":"GIGAZINE","url":"https://gigazine.net/news/20260825-nvidia-avo/","title":"NVIDIAのAIエージェント「AVO」がARC-AGI-3で100％を達成、同じベースモデルの別条件での評価は約30％で「実行基盤(ハーネス)」の重要性が示される"}],"entities":{"anime_titles":[],"manga_titles":[],"work_titles":["ARC-AGI-3"],"studios":["NVIDIA"],"people":[],"type":"news","domain":"other","is_roundup":false},"keyFacts":null}
