OpenAI’s Jalapeño Chip Could Make AI Inference Faster
OpenAI has released the first benchmark results for its Jalapeño inference chip, reporting major gains in performance per watt and response latency. Here is what the results mean for AI infrastructure, agents, and Nvidia.
On this page
OpenAI has published the first measured results from Jalapeño, its custom chip for running trained AI models, and the numbers point to a different way of improving AI performance: make the hardware itself fit the workload. In tests using public models, OpenAI says Jalapeño delivered between 1.5 and 1.9 times more AI work per watt and between 1.7 and 3.6 times lower end-to-end latency than the comparison systems. The more important story is not simply that OpenAI has designed a chip, but that it is trying to control more of the stack that determines how quickly and efficiently AI reaches users.
Jalapeño targets the part of AI people actually use
Jalapeño is an inference chip, meaning it is designed to run an already-trained model and generate responses rather than train a new model from scratch. That distinction matters because inference happens every time someone asks an AI system a question, runs an agent, or sends a request through an application programming interface. OpenAI says Jalapeño is intended to improve both throughput, or how much useful AI work the system can handle, and latency, or how long users wait for a response. OpenAI first unveiled the chip in June as part of a multigeneration hardware platform being developed with Broadcom and other partners.
The first benchmarks show a speed and efficiency advantage
OpenAI tested Jalapeño using InferenceX, a public benchmark from SemiAnalysis that measures the full process of serving an AI request. The comparison covered GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T, giving the chip a test across models from different developers rather than relying only on OpenAI's own models. OpenAI reports that Jalapeño delivered 1.5 to 1.9 times more performance per watt and 1.7 to 3.6 times lower end-to-end latency across the three workloads. On Kimi K2.5 1T specifically, OpenAI reports about 1.5 times higher peak performance per watt and 3.4 times lower end-to-end latency than its comparison system.
Those figures have a practical meaning. A system that produces more useful work from the same amount of electricity can serve more requests without increasing power consumption at the same rate, while lower latency makes interactive applications feel faster. That matters particularly for AI agents, which may perform many model calls in sequence; a small delay repeated across several steps can turn into a noticeably slower task. OpenAI therefore measures Jalapeño not just by peak chip performance, but by the amount of useful work it can deliver while meeting a target user experience.
OpenAI designed the system around AI inference
The architectural difference is that Jalapeño was designed around the way modern language models actually run. When a model processes a prompt, the prefill stage is heavily dependent on computation. During decode, when the response is generated token by token, memory bandwidth becomes a larger constraint. Moving information between processors can add another source of delay. OpenAI says Jalapeño combines the chip, memory, networking, software, and rack-scale system to reduce those bottlenecks instead of treating the accelerator as an isolated component.
One part of that design is keeping model state, including the key-value cache used during generation, closer to the processing resources that need it. Reducing unnecessary movement of that data can reduce communication delays and keep more of the system working instead of waiting. The approach is particularly relevant to agentic workloads, where a model repeatedly reasons, calls a tool, receives information, and generates another response.
The chip is not replacing Nvidia overnight
OpenAI's results are significant, but they do not mean the company has made Nvidia hardware obsolete. The published comparisons are specific benchmark measurements against particular systems and model configurations, while commercial hardware will continue to evolve before Jalapeño reaches broad deployment. OpenAI itself says it will continue widely deploying accelerators from Nvidia and other partners for both training and inference. That makes Jalapeño better understood as an additional source of compute rather than a complete replacement for the existing infrastructure.
There is also a timing issue. OpenAI plans to begin deploying Jalapeño inside its own compute infrastructure by the end of 2026, so the benchmark results describe a system moving toward production rather than a chip already serving OpenAI's entire user base. The company is still completing production qualification, improving the software stack, preparing the system to operate at scale, and validating it across more models.
AI helped build the chip as well as run on it
Perhaps the more unusual part of the project is how AI was used during development. OpenAI says its models helped the team move from initial design to tapeout, the point when the chip design is finalized for manufacturing, in nine months. AI was also used to explore implementations, shorten verification loops, and optimize selected arithmetic circuits. The company designed Jalapeño to be predictable enough that both engineers and AI systems can help program and optimize workloads for it.
OpenAI also reports that Codex working with GPT-Astra brought three open-weight models that were not originally part of the production plan to high performance on Jalapeño within two months. In selected GPT-OSS attention and mixture-of-experts blocks, OpenAI says AI-generated implementations ran 1.5 to 1.8 times faster than existing human-written implementations. Those measurements apply only to selected components, not complete models, so they should not be interpreted as a 1.8-times speedup for an entire AI system.
Why inference efficiency may matter more as agents grow
The economics of AI increasingly depend on what happens after a model has been trained. A more capable model can create demand for more requests, but every request consumes compute, memory, networking capacity, and electricity. If hardware can produce more useful output from the same resources, an AI provider can potentially handle greater demand without increasing infrastructure costs at the same pace. That is the reason Jalapeño's performance-per-watt results may matter as much as its raw speed.
The effect could be even more noticeable for agents. Traditional chat can tolerate some waiting because a user is often making one request at a time. An agent may instead make many model calls while completing a task, making both latency and efficiency cumulative concerns. Hardware optimized for that pattern could allow providers to run more complex agent workflows without simply throwing proportionally more computing power at each task.
Jalapeño is the beginning of a longer hardware strategy
OpenAI describes Jalapeño as the first generation of a multigeneration roadmap rather than a one-off experiment. The company says a second generation is already deep in development and a third is taking shape. That suggests the real test will come later: whether OpenAI can repeatedly improve its own inference hardware while keeping pace with rapidly changing model architectures and competing accelerator designs.
For now, the benchmark results establish a credible starting point, but they remain OpenAI's reported measurements. The company used a public benchmark and published detailed operating points, which makes the claims more useful to examine than an unsupported performance promise, yet independent testing at production scale will still be important. The next meaningful milestone is therefore not another headline benchmark. It is whether Jalapeño can move from controlled tests into a large deployment and deliver the same efficiency and responsiveness across the changing workloads that OpenAI's products actually generate.
Written by


