{"rewrite":{"id":"r_82d616909992bd8055a44361","clusterId":"c_c13b2d7e61d805d47b9a559b","slug":"openai-publishes-benchmark-results-for-custom-inference-chip-jalapeno","model":"deepseek-v4-flash","headline":"OpenAI Publishes Benchmark Results for Custom Inference Chip Jalapeno","summary":"OpenAI held a briefing on Jalapeño, a custom inference chip co-developed with Broadcom, and published benchmark results. The chip reportedly achieves both high throughput and low latency in a single architecture, a combination that typically requires a trade-off. OpenAI plans to deploy the first generation within its internal computing infrastructure by the end of the year.","whyItMatters":"Jalapeño's published benchmarks show it overcoming the usual throughput-versus-latency trade-off in a single architecture, with the advantage growing on larger and more demanding workloads.","webCardHtml":"\u003cp\u003eJalapeño was announced in June 2026 as a custom inference chip optimized for large language models. Since then, OpenAI has run repeated tests on the chip and the systems built around it.\u003c/p\u003e\u003cp\u003eAcross three models, GPT OSS 120B, DeepSeek R1, and Kimi K2.5 1T, peak throughput improved AI processing per 1kW by 1.5 to 1.9 times and cut end-to-end latency by 1.7 to 3.6 times. For highly interactive workloads, the improvement rose to 2.1 to 4.1 times. Peak decode throughput improved across all three models, and throughput per 1kW exceeded previous records.\u003c/p\u003e\u003cp\u003eOpenAI plans to deploy the first generation inside its own infrastructure by the end of the year. A second generation is in development, and a third generation is taking shape.\u003c/p\u003e","blueskyPost":"Jalapeño benchmark results are out. The custom inference chip co-developed with Broadcom shows high throughput and low latency in one architecture. First gen deploys inside OpenAI by year end.","twitterPost":"Jalapeño benchmarks published. OpenAI's inference chip with Broadcom cuts the throughput-latency trade-off: 1.5-1.9x per 1kW, 1.7-3.6x lower latency. First gen deploys internally by year end.","threadsPost":null,"newsletterBlurb":"OpenAI published the first benchmark results for Jalapeño, its custom inference chip co-developed with Broadcom. The chip reportedly achieves high throughput and low latency in a single architecture. The first generation is slated for deployment within OpenAI's own computing infrastructure by the end of the year.","attributionJson":"[{\"source\":\"GIGAZINE\",\"url\":\"https://gigazine.net/news/20260826-jalapeno-ai-inference-results/\",\"title\":\"OpenAI Publishes Benchmark Results for Custom Inference Chip 'Jalapeño', Achieving High Throughput and Low Latency in a Single Architecture\"}]","lintFlagsJson":null,"lintHits":0,"costUsd":0,"inputTokens":4705,"outputTokens":618,"status":"published","repairAttempts":0,"nextRepairAt":null,"factsAttemptedAt":1788009010,"createdAt":"2026-08-29T13:00:51.000Z","publishedAt":"2026-08-29T13:01:45.000Z","updatedAt":"2026-08-29T13:00:51.000Z"},"cluster":{"id":"c_c13b2d7e61d805d47b9a559b","canonicalTitle":"OpenAIがカスタム推論チップ「Jalapeño」のベンチマーク結果を公開、高スループットと低レイテンシを単一アーキテクチャで両立","representativeArticleId":"a_7c60e78d20a38bf26fb1cb9f","sourceCount":1,"writtenSourceCount":1,"writeAttempts":0,"isSolo":true,"entitiesJson":"{\"anime_titles\":[],\"manga_titles\":[],\"work_titles\":[\"Jalapeño\"],\"studios\":[\"Broadcom\"],\"people\":[],\"type\":\"news\",\"domain\":\"other\",\"is_roundup\":false}","contentType":"news","status":"published","firstSeenAt":"2026-08-26T04:45:00.000Z","lastSeenAt":"2026-08-26T04:45:00.000Z","updatedAt":"2026-08-29T13:01:47.000Z"},"attribution":[{"source":"GIGAZINE","url":"https://gigazine.net/news/20260826-jalapeno-ai-inference-results/","title":"OpenAIがカスタム推論チップ「Jalapeño」のベンチマーク結果を公開、高スループットと低レイテンシを単一アーキテクチャで両立"}],"entities":{"anime_titles":[],"manga_titles":[],"work_titles":["Jalapeño"],"studios":["Broadcom"],"people":[],"type":"news","domain":"other","is_roundup":false},"keyFacts":null}
