At the Hot Chips conference on Tuesday, OpenAI unveiled details regarding Jalapeño, featuring the initial batch of benchmark results for the new system. Evaluated against Semianalysis’s InferenceX benchmark, Jalapeño demonstrated superior performance in both tokens per user and throughput per kilowatt compared to current state-of-the-art inference processors.
“The bottom line is that the results show a very, very significant performance advance over state of the art,” stated Richard Ho, OpenAI’s head of hardware, during a press call. “Jalapeño can serve more AI work per unit of power, while also returning responses more quickly. It’s very efficient to serve a lot of customers, but it can also be very low latency.”
While that comparison was made against an Nvidia Blackwell system, OpenAI acknowledges that the competition may have advanced significantly by the time Jalapeño reaches full deployment. Ho estimated that the chip will deploy at the end of 2026 in “very small volumes,” with larger scale deployment expected in 2027.
First announced last October, Jalapeño was developed by OpenAI in close collaboration with Broadcom, with the company’s own models playing a role in the development process. OpenAI intends to position Jalapeño as a multigenerational platform, coordinating the development of AI products, models, chips, and memory.
Due to this comprehensive full-stack approach, OpenAI was able to tackle specific phases in the inference process that typically cause friction. Specifically, Jalapeño is designed to reduce delays during the prefill and communication phases of processing, which the company identifies as common bottlenecks.
“We designed Jalapeño to minimize data movement and communication delays,” the company explained in a blog post regarding the results. “This ensures that model state, including the KV cache used for generating responses, can be explicitly placed and kept local while the system activates the optimal combination of compute, memory, and networking for each inference phase.”
The Editorial Staff at AIChief is a team of professional content writers with extensive experience in AI and marketing. Founded in 2025, AIChief has quickly grown into the largest free AI resource hub in the industry.
