Broadcom CEO Hock Tan handed the chip to OpenAI CEO Sam Altman.
SAN FRANCISCO: On June 24 2026, OpenAI and Broadcom delivered the first engineering samples of Jalapeño OpenAI’s first proprietary AI accelerator at OpenAI’s San Francisco headquarters. Broadcom CEO Hock Tan and Semiconductor Solutions President Charlie Kawwas handed the chip to OpenAI CEO Sam Altman and President Greg Brockman, marking the first physical milestone of a custom silicon partnership the two companies first disclosed in October 2025. The chip is designed from scratch for large language model inference and is already running GPT-5.3-Codex-Spark, a real-time coding model at production-target frequency and power in the lab. Initial deployment is targeted for late 2026 with a full production ramp in 2027 and gigawatt-scale data centre deployment alongside Microsoft and other partners through 2029.
What Jalapeño Is and What It Does
To understand why this matters the distinction between a GPU and an ASIC is essential. A graphics processing unit is a general-purpose parallel processor. It handles a wide variety of computing tasks effectively, which also means it carries architectural overhead that is unnecessary for a specific repetitive workload like AI inference. An Application-Specific Integrated Circuit is designed from the ground up for one job. Jalapeño is an ASIC built solely to run pre-trained AI models in response to user queries the process called inference.
Inference is what happens every time someone sends a message to ChatGPT. Training a model is a one-time capital expenditure. Inference is a continuous operational cost that scales directly with every user query. As models grow larger and user bases expand into hundreds of millions, the cost per token the unit economics of running an AI product becomes the central financial challenge of the business.
Jalapeño addresses this by minimising the architectural inefficiencies that inflate inference costs on general-purpose GPUs. The chip reduces data movement between processing units and memory, balances compute, memory and networking resources on the die and integrates Broadcom’s Tomahawk networking silicon for low-latency communication across multi-chip clusters. Richard Ho, who leads OpenAI’s hardware program and previously led parts of Google’s TPU programme explained the design intent directly: “We optimized the architecture around the kernels, memory movement, networking, and serving patterns that matter most for frontier AI models. Based on early testing, Jalapeño will efficiently execute our most important workloads close to the hardware’s theoretical limits.”
The development also used OpenAI’s own models to accelerate parts of the chip design process a self-referential loop where the models being served are helping build the infrastructure that will serve future models. OpenAI described the nine-month timeline from design to tape-out as potentially the fastest ASIC development cycle ever achieved in high-performance advanced semiconductors. The chip was fabricated by TSMC with Celestica handling board design, rack assembly and system integration.
The Performance
Hock Tan told Bloomberg that Jalapeño is showing cost savings of roughly 50 percent per inference token compared to current-generation GPUs and told Reuters that performance is on par with Nvidia Blackwell chips and Google TPUs. OpenAI’s own official announcement is deliberately more measured, stating only that early testing shows performance per watt substantially better than current state-of-the-art with final numbers still being measured and a detailed technical report promised in the coming months.
No independent third-party benchmarks exist at the time of announcement. The engineering samples are running inside OpenAI’s closed laboratories. What is confirmed is that the chip is operating at production-target frequency and power without failure a meaningful hardware milestone that indicates architectural stability and manufacturing readiness. Successfully running GPT-5.3-Codex-Spark, a latency-first model optimised to generate code at over 1,000 tokens per second suggests the chip’s memory pipelines and interconnects can handle highly interactive, real-time workloads. The full picture will arrive with the technical report.
Why OpenAI Needed to Build This
The financial imperative behind Jalapeño is clear from OpenAI’s 2025 audited financials. The company generated USD 13.07 billion in revenue but incurred USD 34 billion in total operational expenses, resulting in an operating loss of nearly USD 20.92 billion. Compute infrastructure drove the bulk of that outflow research, development and compute costs totalled USD 19.18 billion, including USD 10.59 billion paid to Microsoft for computing infrastructure alone.
Reducing inference costs directly improves operating margins as user queries scale. Greg Brockman framed the strategic logic in the official announcement: “By designing more of the stack ourselves, we can serve more intelligence with greater efficiency and keep pushing advanced AI toward broader access.” This mirrors what Google has done with TPUs, Amazon with Trainium and Inferentia and Microsoft with its Maia 200 accelerator. OpenAI is joining a competitive field where controlling the full infrastructure stack chip architecture, kernels, memory systems, networking and deployment is increasingly seen as necessary for profitability at scale.
Jalapeño is not a replacement for Nvidia GPUs across OpenAI’s operations. Pre-training and large foundational model runs will continue to rely on Nvidia’s flexible, high-bandwidth GPU clusters. Jalapeño is a specialised tool for the inference layer the high-volume continuous-cost workload where ASIC efficiency advantages are most commercially meaningful. Nvidia’s CUDA software ecosystem built over decades and embedded in the workflows of millions of developers remains structurally secure in the near term.
The Broader Landscape
The custom chip trend extends well beyond OpenAI and Broadcom. Alibaba’s semiconductor arm T-Head unveiled the Zhenwu M890 in May 2026, featuring 144 gigabytes of high-bandwidth memory optimised for autonomous AI agents. Huawei’s Ascend 950DT, co-designed with researchers from DeepSeek reportedly cut DeepSeek’s API serving costs by 75 percent. Qualcomm is separately reported to be in negotiations with ByteDance to provide custom inference chip designs though neither company has officially confirmed this. The structural shift toward purpose-built inference silicon is accelerating across both Western and Chinese technology ecosystems.
For India, the implications run in two directions. Indian IT companies like TCS and Infosys building agentic AI workflows on OpenAI’s API, would see deployment costs fall significantly if inference prices drop strengthening their competitive positioning in AI-led services. Conversely Indian data centre operators who have committed to Nvidia GPU infrastructure face a longer-term question about how the economics of general-purpose GPU rental evolve as purpose-built inference ASICs take an increasing share of production workloads. India’s domestic semiconductor push under ISM 2.0 and the C-DAC AUM processor programme are moving in a similar direction, though the execution timeline for domestic AI silicon remains considerably longer than the nine months it took OpenAI and Broadcom to bring Jalapeño from design to working samples.
The technical report OpenAI has promised will be the real test of what Jalapeño can do. Until then the 50 percent cost saving figure belongs to Broadcom’s CEO, not to a benchmark.
