OpenAI has announced the Jalapeño inference chip, which the company said outperforms Nvidia's GB300 chip in power efficiency and response speed.

This development marks a significant shift as major AI software developers move toward designing their own custom hardware. By reducing reliance on external chip providers, OpenAI aims to lower operational costs and optimize the hardware specifically for large language models.

The announcement took place at the Hot Chips conference, where OpenAI detailed the collaboration with Broadcom Inc. to build the hardware. According to internal tests conducted by OpenAI, the Jalapeño chip delivers more AI work per unit of power and provides faster inference compared to the GB300 [1, 2, 3].

OpenAI said the design-to-tape-out development cycle for the chip lasted approximately 16 months [4, 5]. This rapid turnaround suggests an aggressive timeline to scale internal infrastructure to meet the growing demands of generative AI.

While some early reports mentioned the Blackwell architecture, multiple sources corroborated that the specific benchmark target was the GB300 [1, 2, 3]. The chip was designed to address the bottlenecks associated with power consumption and latency in high-scale AI deployments.

The partnership with Broadcom provided the necessary engineering expertise to move from conceptual design to a physical chip. OpenAI said the goal was to create a specialized environment where the hardware and software are co-optimized for maximum performance [1, 2].

The Jalapeño chip delivers more AI work per unit of power and provides faster inference.

The move toward proprietary silicon indicates that the 'AI arms race' is shifting from software capabilities to hardware efficiency. If OpenAI can successfully deploy the Jalapeño chip at scale, it reduces the industry's systemic dependence on Nvidia's supply chain and creates a vertical integration model where the model creator controls the entire stack from the chip to the user interface.