OpenAI unveils Jalapeño AI chip, claims 1.9x efficiency over Nvidia
Serge Bulaev
OpenAI introduced its Jalapeño AI chip, which they claim may be 1.5 to 1.9 times more efficient than Nvidia's latest chips, based on their own early tests. The chip was designed in about nine months using AI-assisted tools, which might have made the process faster than usual. Jalapeño uses less power (700 W) compared to similar Nvidia chips and is made for agentic and reasoning workloads. These numbers come from OpenAI and have not yet been confirmed by independent labs, so the actual performance may change as more data becomes available.

OpenAI has unveiled its first custom Jalapeño AI chip, an accelerator it claims delivers up to 1.9x more efficiency than Nvidia's latest hardware. Presented at Hot Chips 2026, the chip was developed in a remarkably fast nine-month cycle using AI-assisted design tools. Jalapeño is an inference-focused ASIC (Application-Specific Integrated Circuit) designed to power agentic and reasoning workloads with significantly lower power consumption.
What is Jalapeño and why did OpenAI build it?
Jalapeño is OpenAI's first custom-designed AI accelerator chip, created for internal use to power inference and reasoning workloads. Built in partnership with Broadcom on an advanced TSMC process, it represents OpenAI's strategic move to develop specialized hardware optimized for its own advanced AI models.
Jalapeño is the first entry in OpenAI's "Intelligence Processor" family, fabricated on an advanced TSMC process node. This project marks a strategic shift for OpenAI, moving beyond reliance on general-purpose GPUs to create purpose-built silicon tailored to its unique model architectures. Hardware leads Richard Ho, Ravi Narayanaswami, and Chris Leary presented the chip, identifying Broadcom as the silicon implementation partner responsible for the chip's production and system integration. According to OpenAI, designing custom hardware enables them to embed insights from frontier model development directly into silicon, aiming to "unlock new levels of capability and intelligence."
How does Jalapeño compare to Nvidia's Blackwell chips?
In its initial benchmarks, OpenAI claims Jalapeño delivers 1.5× to 1.9× higher throughput per watt and 1.7× to 3.6× lower end-to-end latency than Nvidia's GB200 and GB300 systems. These tests were conducted using large models like GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T.
A key differentiator is power consumption: the Jalapeño test part has a 700W thermal design power (TDP), contrasting sharply with the ~1,200W for a GB200 and ~1,400W for a GB300 in the comparison setups. It is critical to note that these figures are from OpenAI's own disclosure of early engineering samples and have not yet been independently validated. Both Tom's Hardware and SemiAnalysis position these results as a significant performance-per-watt advantage for inference, not as a direct replacement for general-purpose training GPUs.
What makes the nine-month development timeline so remarkable?
The nine-month cycle from initial design (RTL) to tapeout is an extraordinary achievement in high-performance computing. Custom ASIC projects of this complexity typically take over 24 months, making Jalapeño's development timeline potentially one of the fastest ever for an advanced semiconductor.
This acceleration was made possible by leveraging advanced design tools and methodologies. According to industry reports, AI-assisted Electronic Design Automation (EDA) flows are increasingly helping with design-space exploration, placement and routing, and verification, allowing engineering teams to make late-stage changes close to the design freeze. As noted in an EDN analysis of similar timelines, this reflects a significant acceleration from AI tools, though human engineers still retain final signoff and control over the process.
What workloads is Jalapeño optimized for?
Jalapeño is a specialized accelerator explicitly designed for AI inference, particularly for agentic and reasoning workloads, not large-scale model training. OpenAI identified a need for a purpose-built chip focused on key metrics for serving models to users:
- Low-latency inference across multiple chips
- Maximum performance-per-watt efficiency
- High memory bandwidth
- Reduced cost per token served
This focus aligns with a wider industry trend where hardware is increasingly tuned for specific inference patterns as AI models become more complex. The chip is being tested with OpenAI's software infrastructure, with first silicon validation completed in early 2026.
Will OpenAI sell Jalapeño chips?
No, OpenAI has confirmed that Jalapeño is an internal project for its own infrastructure and will not be sold as a merchant silicon product. This is part of a larger strategic collaboration with Broadcom, first announced in 2025, to deploy up to 10 gigawatts of custom AI accelerators.
The goal is to reduce token-level serving costs and power consumption across OpenAI's services. Initial deployment is targeted for the second half of 2026, with completion targeted by end of 2029. This indicates that Jalapeño ASICs will coexist with and complement OpenAI's existing large-scale GPU fleets rather than replace them immediately.
Jalapeño's Key Metrics (OpenAI-Reported)
- Efficiency: 1.5× - 1.9× greater throughput per watt vs. Nvidia GB200/GB300
- Latency: 1.7× - 3.6× lower end-to-end latency on tested models
- Power: 700W TDP, compared to 1,200W - 1,400W for Nvidia systems
- Development: Nine months from RTL kick-off to tapeout on advanced TSMC process
These metrics, while impressive, are preliminary and may be updated as the silicon matures and undergoes third-party testing.