
AI NEWS
OpenAI’s Jalapeño chip is built for fast inference at scale, benchmarks show
OpenAI unveiled Jalapeño, a custom AI inference chip developed with Broadcom at the Hot Chips conference. Benchmarks indicate it outperforms current Nvidia Blackwell systems in throughput and tokens per user while maintaining low latency. The chip targets high efficiency for serving many customers simultaneously. Deployment is scheduled to begin in small volumes by late 2026, expanding significantly in 2027.
THE NEWS
What happened
OpenAI unveiled Jalapeño, a custom AI inference chip developed with Broadcom at the Hot Chips conference. Benchmarks indicate it outperforms current Nvidia Blackwell systems in throughput and tokens per user while maintaining low latency. The chip targets high efficiency for serving many customers simultaneously. Deployment is scheduled to begin in small volumes by late 2026, expanding significantly in 2027.
CONTEXT
Why it matters
OpenAI has revealed Jalapeño, a new custom AI inference chip developed in partnership with Broadcom. Initial benchmark results from the Hot Chips conference show significant performance gains over current state-of-the-art processors like Nvidia Blackwell. The chip delivers higher throughput per kilowatt and more tokens per user while maintaining low latency. OpenAI plans to deploy Jalapeño in small volumes by the end of 2026, with a major rollout expected in 2027. Designed to minimize communication delays and data movement, this full-stack approach aims to serve more customers efficiently. This development marks a pivotal shift in AI hardware infrastructure, potentially reshaping how large-scale inference is handled across North American tech sectors.
AT A GLANCE
Key facts
- Jalapeño was announced at the Hot Chips conference with initial benchmark results released.
- The chip demonstrated higher tokens per user and throughput per kilowatt than Nvidia Blackwell systems on SemiAnalysis' InferenceX benchmark.
- OpenAI's head of hardware, Richard Ho, described the performance advance as very significant.
- Development involved close collaboration between OpenAI and Broadcom, utilizing OpenAI models to assist in chip design.
- Deployment is planned for late 2026 in small volumes with major expansion expected in 2027.
- The architecture minimizes data movement and communication delays during prefill and inference phases.
- Jalapeño allows model state, including KV cache, to remain local while optimizing compute and networking resources.
SOURCE
Original source
This article is based on information published by TechCrunch AI.



