OpenAI, together with semiconductor giant Broadcom, has unveiled its first AI chip, "Jalapeño"[1]. It is a dedicated chip built specifically for the "inference" workloads that answer user queries on services such as ChatGPT, and the companies say they took it from design to manufacturing in just nine months[1][2]. Early testing suggests its performance per watt will substantially exceed the current state of the art, and full-scale deployment in data centers is planned to begin at the end of 2026[1].

What Kind of Chip Is Jalapeño

Jalapeño is an AI accelerator (a chip that specializes in speeding up AI computation) that OpenAI calls an "Intelligence Processor"[1]. Inference here refers to the process in which a trained AI model assembles an answer to a user's question, such as when ChatGPT generates text[3].

A key point is that it is not a repurposed existing GPU but was designed from a blank slate for the inference of large language models (LLMs)[1]. OpenAI set the chip's architecture based on its knowledge of its models, kernels, and serving systems, aiming to reduce data movement and balance compute, memory, and networking resources to reach realized performance close to the theoretical peak[1]. The chip was revealed when Broadcom's leaders handed it to OpenAI CEO Sam Altman and President Greg Brockman[1].

Nine-Month Development and the Performance Outlook

OpenAI says it completed the work from the start of design to manufacturing tape-out (finalizing the production design) in nine months, positioning this as one of the fastest ASIC (application-specific integrated circuit) development cycles in high-performance semiconductors[1][2]. It used its own AI models for part of the design and optimization, noting that having AI help with chip design can shorten the development cycle[1].

On performance, engineering samples placed in the lab are already running real machine-learning workloads at the frequency and power targeted for production[1]. Models including "GPT-5.3-Codex-Spark" are said to be running on them[1]. As an early-test result, the company states that performance per watt is expected to substantially exceed the current state of the art[1]. However, final performance is still being measured, and a detailed technical report is planned for release within the coming months, so the current figures remain OpenAI's own account[1].

On manufacturing, Broadcom handles silicon implementation and networking technologies (including communication chips called Tomahawk), while Celestica takes charge of boards, racks, and overall system integration[1][2].

A Full-Stack Strategy Supporting ChatGPT

OpenAI describes the chip as part of a "full-stack" strategy in which it designs in-house even the infrastructure beneath its models and products[1]. By optimizing everything from chip architecture to kernels, memory, networking, task scheduling, and deployment systems toward the same goal, it aims to run its own models faster, more cheaply, and more reliably[1].

Inference is the point where AI reaches users, and the company frames improvements in cost, speed, and reliability as translating directly into faster ChatGPT responses and lower costs for developers using the API[1]. Jalapeño is described as the first step in a multi-generation compute platform, with initial deployment at the end of 2026 and plans to expand at gigawatt scale alongside data center partners including Microsoft[1][2].

A Counterweight to NVIDIA's Dominance

Major media outlets read the announcement as a move to ease dependence on NVIDIA, which holds an overwhelming share of AI semiconductors[3]. As large players such as Google and Amazon push ahead with their own chips, OpenAI has joined the trend of holding its own dedicated silicon to diversify compute procurement and contain costs[3]. That said, Jalapeño has yet to be mass-produced and prove itself in real operation, and whether its claimed performance advantage holds up in the field remains to be verified[1][3].

Summary

Jalapeño, unveiled by OpenAI with Broadcom, is the company's first in-house AI chip specialized for LLM inference. It went from design to manufacturing in nine months and is said to exceed the current state of the art in performance per watt, though final performance is still being measured and a detailed report is months away. Broadcom and Celestica support production, and deployment at gigawatt scale with partners such as Microsoft is planned from the end of 2026. It stands as a symbol of OpenAI's full-stack strategy of designing its own infrastructure, and how far it will reshape the NVIDIA-dominated landscape of AI semiconductors is the question going forward.

Source:https://openai.com/index/openai-broadcom-jalapeno-inference-chip

Source:https://investors.broadcom.com/news-releases/news-release-details/openai-and-broadcom-unveil-llm-optimized-intelligence-processor

Source:https://www.cnbc.com/2026/06/24/openai-and-broadcom-reveal-jalapeno-first-ai-chip-in-partnership.html