OpenAI and Broadcom have unveiled Jalapeño, OpenAI’s first custom “Intelligence Processor”, marking a major step in the company’s push to build more of the AI infrastructure stack behind ChatGPT, Codex, its API products and future agentic systems.
The chip has been designed specifically for large language model inference — the stage where trained AI models generate responses for users — rather than being adapted from general-purpose AI accelerators.
OpenAI said early testing points to performance-per-watt “substantially better” than current state-of-the-art systems, although final benchmark details are expected in a technical report in the coming months.
Built for faster, cheaper AI responses
Jalapeño has been architected around the workloads OpenAI runs every day, including model kernels, memory movement, networking and serving patterns.
The aim is to combine high throughput with lower latency, helping make AI products faster, more reliable and more affordable at scale. In practical terms, OpenAI says infrastructure gains could translate into quicker ChatGPT responses, longer-running Codex tasks, cheaper API products and more dependable access during periods of heavy demand.
Nine-month development cycle
OpenAI said the chip moved from initial design to manufacturing tape-out in nine months, helped by close software-hardware co-development with Broadcom and the use of OpenAI models in parts of the design and optimisation process.
Engineering samples are already running machine learning workloads in the lab at production target frequency and power, including GPT-5.3-Codex-Spark.
Broadcom and Celestica (TSX:CLS) support scale-up
The Jalapeño platform is being developed with Broadcom and Celestica (TSX:CLS), with Broadcom contributing silicon implementation, networking and connectivity technologies, including Tomahawk networking silicon.
Celestica is supporting board, rack and system integration as OpenAI prepares the platform for large-scale deployment.
Gigawatt-scale roadmap
Jalapeño is the first chip in what OpenAI and Broadcom describe as a multi-generation compute platform, with initial deployment targeted by the end of 2026 and expansion planned in the years that follow.
Broadcom CEO Hock Tan said the collaboration was aimed at enabling “gigawatt scale data centers with Microsoft and other partners beginning in 2026”.
Why it matters
The announcement underscores how central compute infrastructure has become to the AI race. By designing its own inference hardware, OpenAI is moving deeper into the full stack — from products and models to chips, networking and deployment systems.
For users and businesses, the promise is simple: more efficient inference could help make advanced AI faster, more available and less expensive to run.