Benchmarks verified by SemiAnalysis show that the chip currently outperforms flagship products from Nvidia, AMD, and Google when running several prominent open-source models.
The success of Jalapeño is significant because it achieves industry-leading efficiency in "performance per watt," a critical metric for companies limited by data center power capacity.
While competitors like Nvidia often rely on complex software techniques to boost speed, OpenAI’s chip delivered superior results using "single-token prediction," or generating one piece of data at a time without those advanced shortcuts.
This efficiency allows OpenAI to produce more model responses within the strict energy limits of modern infrastructure, potentially reducing the high costs associated with operating large-scale AI services.
OpenAI achieved this rapid development—moving from initial hiring to a finished design in roughly 16 months—by using its own AI models to assist in the chip’s design and kernel programming.
The hardware features "HBM4," a next-generation high-bandwidth memory that enables faster data access than the memory found in most current chips.
While the chip is currently in a testing phase, OpenAI plans to gradually ramp up production throughout 2027, eventually deploying the hardware across data centers to power its future models and enterprise services.