OpenAI reveals benchmark results for its Jalapeño inference chip
OpenAI presented the first benchmark results for its new Jalapeño chip at the Hot Chips conference, showing it outperforms current state-of-the-art inference hardware in both speed and power efficiency. The chip, developed with Broadcom, is expected to begin limited deployment by the end of 2026.

At the Hot Chips conference on Tuesday, OpenAI gave a more detailed look at its new inference chip, Jalapeño, sharing the first round of independent benchmark results. Tested on Semianalysis's InferenceX benchmark, the chip delivered both a higher number of tokens processed per user and greater throughput per kilowatt than the best inference processors currently available on the market.
Richard Ho, OpenAI's head of hardware, told reporters on a press call that the results point to a very significant performance advance over existing technology. He explained that Jalapeño can handle more AI workload per unit of power while also responding faster, making it well suited both for serving large numbers of customers efficiently and for applications requiring very low latency.
The comparison was made against an Nvidia Blackwell system, though Ho acknowledged that competitors may advance considerably before Jalapeño reaches full-scale deployment. He estimated that the chip would begin shipping in very small volumes by the end of 2026, with broader rollout planned for 2027.
Built with Broadcom
Jalapeño was first announced last October and has been developed by OpenAI in close collaboration with Broadcom, with OpenAI's own models assisting in the design process. The company intends to turn Jalapeño into a multigenerational platform, developing AI products, models, chips and memory together as an integrated system.
This full-stack approach allowed OpenAI to target specific stages of the inference process that commonly cause slowdowns. Jalapeño is designed to reduce delays during the prefill and communication phases, which the company identifies as frequent bottlenecks. According to OpenAI's blog post detailing the results, the system minimizes data movement and communication delays, allowing model state — including the cache used while generating responses — to be kept close to the compute, memory and networking resources needed at each stage of processing.


