OpenAI and Broadcom unveil Jalapeño, OpenAI's first custom LLM inference chip

openaibroadcomjalapenoinference-chipcustom-siliconintelligence-processor+12high-risk claims
Screenshot of the OpenAI 'OpenAI and Broadcom unveil LLM-optimized inference chip' announcement page archived by the Wayback Machine on 2026-06-24 at 20:14:46 UTC, showing the headline 'Designed to be the best inference platform for LLMs' above the section 'Nine-month tape-out, accelerated by OpenAI models'.
Screenshot: OpenAI / Broadcom announcement (openai.com, 2026-06-24) — captured 2026-06-25 from the Wayback Machine archive snapshot 20260624201446 (the live openai.com URL returned HTTP 403 to a non-browser Cloudflare challenge). Credit: OpenAI. License: OpenAI published material, used for editorial commentary under fair use.

On 2026-06-24, OpenAI and Broadcom (NASDAQ: AVGO) jointly unveiled Jalapeño, OpenAI’s first custom Intelligence Processor — a blank-slate LLM inference accelerator co-developed with Broadcom and Celestica, taken from initial design to tape-out in nine months, with engineering samples already running ML workloads in the lab at production target frequency and power. The first chips were physically delivered to OpenAI CEO Sam Altman and President Greg Brockman by Broadcom CEO Hock Tan and President Charlie Kawwas (OpenAI, 2026-06-24; Wayback Machine archive, 2026-06-24).

OpenAI is the first US frontier AI lab to publicly attach its name to a custom inference silicon program. The performance numbers — “substantially better” per watt and the 9-month tape-out — are vendor claims, not independent measurements. A detailed technical report is promised “in the coming months.”

What was announced

Load-bearing quotes

“Jalapeño is part of our long-term full-stack infrastructure strategy to make compute more abundant, resulting in AI which is faster, more reliable, more affordable for people and businesses, and can be used to solve more important problems.” — Greg Brockman, President and Co-Founder, OpenAI (OpenAI, 2026-06-24)

“Jalapeño was designed from the ground up for LLM inference. We optimized the architecture around the kernels, memory movement, networking, and serving patterns that matter most for frontier AI models.” — Richard Ho, OpenAI hardware lead (OpenAI, 2026-06-24)

Why it matters

1. OpenAI is the first US pure-model-lab to ship custom inference silicon. Google has TPU (since 2015), Amazon has Trainium and Inferentia, but OpenAI is the first to publicly attach its name to a custom-silicon program. The comparison set for any other frontier lab considering in-house inference silicon moves from “impossible” to “table stakes.”

2. Vertical integration follows the Google TPU / Pathways pattern. Chip architecture, kernels, memory systems, networking, scheduling, deployment, and product now sit in the same organization. Any frontier lab buying 100% of inference compute from NVIDIA is now the outlier.

3. The 9-month tape-out is the most quotable, most contestable claim. If independently confirmed, it would be the fastest published ASIC development cycle in high-performance semiconductors. The claim is Broadcom’s, prefixed with “we believe.” The technical report is the load-bearing source.

4. Competitive pressure on NVIDIA, AMD, and Google TPU. A frontier lab buying inference silicon from anywhere but NVIDIA changes the supplier-mix math. No benchmark exists today to size the shift.

Practical implications

Risks and caveats

  1. “Substantially better” is a vendor claim; no benchmark, no comparison to NVIDIA or Google TPU.
  2. The 9-month tape-out is also a vendor claim, prefixed with “we believe.”
  3. “GPT-5.3-Codex-Spark” is a workload, not a public product — no release date, no API, no pricing.
  4. “Gigawatt scale” is unspecific: no target, no timeline beyond “beginning in 2026,” no per-region allocation.
  5. Microsoft is named without specifics; Microsoft has not publicly named Jalapeño.
  6. No shipping date, no SKU, no price. Engineering samples in the lab, running internal workloads.
  7. No independent verification: no third-party benchmark, no analyst note, no research-lab measurement.

What to watch

Sources