Home / News / OpenAI's Jalapeño Chip Cuts AI Latency Up to 3.6x in First Benchmarks
OpenAI

OpenAI's Jalapeño Chip Cuts AI Latency Up to 3.6x in First Benchmarks

Aug 26, 20265 min read
OpenAI's Jalapeño Chip Cuts AI Latency Up to 3.6x in First Benchmarks

News Summary

OpenAI has published the first independent benchmark results for Jalapeño, the custom AI inference chip it co-developed with Broadcom, showing the silicon completing real-world language model workloads faster and with less power than comparable commercial GPU setups. The figures, released on August 25, 2026 at 3:00 PM Eastern Time, mark the first time OpenAI has backed up its earlier performance claims about the chip with third-party testing data rather than internal projections.

What the benchmarks measured

OpenAI ran Jalapeño through the SemiAnalysis InferenceX benchmark suite, an independent testing framework widely used to compare AI accelerators on realistic inference workloads rather than synthetic peak-throughput numbers. The tests covered three open-weight language models chosen to represent different scales and use cases: GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T.

Rather than isolating a single metric, the suite measures how a chip performs across the full inference pipeline, including the prefill phase where a model processes an incoming prompt and the token-generation phase where it produces a response. OpenAI has said this dual focus reflects how the chip was actually engineered: to reduce the data-movement and communication delays that tend to bottleneck real deployments, not just to post a high number on a single leaderboard metric.

Headline performance figures

According to the published results, Jalapeño's end-to-end response latency came in 1.7 to 3.6 times faster than standard commercial GPU deployments running the same models. On power efficiency, the chip completed 1.5 to 1.9 times more total work per watt of electricity at peak capacity. For interactive, low-latency workloads such as AI agents that need to execute many rapid steps in sequence, OpenAI reported gains of 2.1 to 4.1 times over comparable setups.

In a more specific head-to-head comparison, OpenAI and Broadcom said Jalapeño, running at a rated thermal design power of 700 watts, outpaced Nvidia's GB300 flagship GPU, which is rated at roughly double that power draw. On that comparison, Jalapeño delivered up to 1.9 times more throughput per kilowatt of electricity consumed and as much as 3.6 times lower latency. Notably, OpenAI said the chip sustained real-world power draw under 550 watts during testing, below its rated ceiling.

Why OpenAI built its own chip

Jalapeño is OpenAI's first custom silicon and was designed from the outset as a purpose-built inference accelerator rather than a repurposed training chip or a general-purpose processor. OpenAI and Broadcom first unveiled the chip on June 24, 2026 Eastern Time, saying it moved from initial design to manufacturing tape-out in about nine months, a pace the companies described as among the fastest development cycles achieved for a chip of this complexity. At that time, Broadcom's leadership also pointed to substantially lower operating costs compared with typical general-purpose GPUs used for inference.

The rationale behind the project is straightforward: as AI companies scale up chatbot and agent products that respond to millions of users in real time, the cost and energy consumed by inference, running an already-trained model to answer queries, has become a larger share of total spending than the one-time cost of training the model. A chip tuned specifically for inference workloads, rather than a general-purpose accelerator that also has to handle training, can in principle be more efficient at the specific job it needs to do.

What comes next

OpenAI has indicated that Jalapeño remains in an early testing phase and that broader deployment across its infrastructure is planned for 2027. The company has not yet disclosed which specific products or services will run on the chip first, nor detailed manufacturing volumes. For now, the newly published benchmark data serves mainly as validation that the chip's real-world performance, tested independently rather than described only in OpenAI's own marketing materials, lines up with the efficiency and latency goals the company set when the project was first announced.

The broader significance for the AI industry is that OpenAI joins a small group of major AI developers, alongside companies such as Google and Amazon, that have moved to design their own inference silicon rather than relying solely on off-the-shelf GPUs from merchant chipmakers. Analysts tracking the space have noted that if Jalapeño's efficiency gains hold up at scale, it could meaningfully lower the cost of running large language models for millions of concurrent users, a cost that has become one of the central constraints on how widely AI assistants and agents can be deployed.

OpenAIAI chip