OpenAI is moving further into AI hardware with Jalapeño, its first custom-designed inference chip, developed with Broadcom. The company aims to make AI responses faster and improve power efficiency as demand for computing continues to grow. OpenAI plans to begin deploying the chip by the end of 2026.
The company says it is also working on later versions of the processor. Jalapeño will not replace NVIDIA hardware across OpenAI’s infrastructure. Instead, it will provide specialized computing capacity for AI inference.
Jalapeño is an application-specific integrated circuit, or ASIC, built for the stage when trained AI models generate responses. OpenAI designed the processor around the computing, memory and networking needs of its own AI systems.
Inference involves two main stages. The first is prefill, where the system processes a user’s prompt. The second is decode, where the model generates a response one token at a time.
OpenAI says Jalapeño addresses the demands of both stages. Its design also aims to keep important model information, including the KV cache, close to the processing hardware. This can reduce the amount of data that needs to move between processors and memory.
Richard Ho, OpenAI’s hardware chief, said Jalapeño is intended to combine high throughput with low latency. OpenAI said the first chip design reached manufacturing tape-out within nine months.
OpenAI tested Jalapeño against leading Nvidia systems using GPT-OSS 120B, DeepSeek R1 670B and Kimi K2.5 1T. The company reported 1.5 to 1.9 times more AI work per watt at peak throughput across the tested workloads.
The tests also showed 1.7 to 3.6 times lower end-to-end latency, according to OpenAI. The company reported 2.1 to 4.1 times higher performance for highly interactive workloads.
OpenAI said that the results come from its own benchmarks. The figures depend on the testing methods and the NVIDIA systems used for comparison.
The development gives OpenAI greater control over hardware designed around its models and software. It also reflects the company’s effort to improve the speed and energy efficiency of AI services as inference workloads expand.
Also Read: Samsung Sees No End to AI Chip Shortage Until 2028; Here's What Fueling Demand