
OpenAI's Jalapeño chip outperforms Nvidia Blackwell in inference benchmarks
SemiAnalysis testing shows OpenAI's self-designed Jalapeño inference chip beats Nvidia's Blackwell on performance per watt — with single-token prediction and no speculative decoding.
OpenAI has built its own inference chip, and the first benchmarks suggest it can compete with hardware from Nvidia, AMD and Google. The results, published on August 25, come from SemiAnalysis running its InferenceX suite inside OpenAI's labs together with the company's engineers.
What Jalapeño is
The chip, code-named Jalapeño, was announced at the Hot Chips conference and is being developed in partnership with Broadcom. Design work began in the middle of 2024 and went from initial hiring to manufacturing tape-out in roughly 16 months, with the CoWoS tape-out completed in November 2025. The published results were gathered on the A0 stepping, about nine months into the program; a B0 stepping is already in the fab and is said to improve performance per watt by around 25 percent.
Performance per watt
The headline figure is token throughput per all-in utility megawatt. On that measure Jalapeño beats Blackwell in almost every scenario tested, and it does so with single-token prediction only — no speculative decoding and no prefill-decode disaggregation — while the comparison systems ran in their best configurations with multi-token prediction. At low concurrency the chip passed 700 tokens per second per user on DeepSeek R1; on Kimi K2.5 it approached 700 tokens/s/user and was more than nine times faster than the next best chip at 100 tokens/s/user. GSM8k evaluations came out on par with Nvidia parts.
A generalized chip and 2,048-accelerator racks
Against the idea that such chips are tuned only for their owner's models, SemiAnalysis describes Jalapeño as a generalized inference accelerator: it ran several open models and the InferenceX suite, and OpenAI even showed Doom running on it, ported with Codex prompts. The chip uses a reticle-sized compute die manufactured on TSMC's N3P, an I/O chiplet on N3E, PCIe Gen 5 links to an x86 host CPU, and HBM4 memory with about 15.4TB/s per package. A rack holds 128 chips connected through 102.4Tb/s Tomahawk 6 switches, and 16 racks — 2,048 accelerators — form one scale-up domain.
Caveats and what comes next
The publication flags limits: all numbers were supplied by OpenAI, the runs it verified in person covered an 8k1k workload, and no AgentX results exist yet. The authors also argue that comparing Jalapeño to Blackwell is not entirely fair, since Jalapeño uses HBM4 — the same generation as Rubin. On cost per token, Jalapeño and Vera Rubin come out roughly level, but Jalapeño's figures are achieved without speculative decoding, which typically cuts cost per token several-fold. Production is scheduled to ramp through 2027, with most output expected at the end of next year and a 100MW deployment as the next goal.
SiTech — AI-powered web development
We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.