OpenAI Jalapeño: Better Than Nvidia Blackwell
newsletter.semianalysis.com · 4,903 words · saved by 2 readers
OpenAI’s self-designed ASIC compared with Rubin, Jalapeño’s TCO, throughput per MW, and spicy deets
OpenAI has spent the past couple years quietly building “Jalapeño,” an inference chip just announced at Hot Chips. Rumors of a successful tapeout had been swirling for a while. But now we have details. OpenAI invited us to look at their chip, go to their labs to check out how real it is, and benchmark it with our InferenceX suite. In June, OpenAI unveiled the chip program in partnership with Broadcom, built from a blank slate exclusively for LLM inference. Design work began in the middle of 2024, going from initial team hiring to manufacturing tape-out in ~16 months, an extremely fast ASIC…
saved by
related reading
- An Interview with MatX CEO Reiner Pope About LLM Chipschipstrat.com
- AI Chip Architecturesjepeake.com
- Taalas is what Etched should have been.zach.be
- The Short Case for Nvidia Stock | YouTube Transcript Optimizeryoutubetranscriptoptimizer.com
- TPU Inference Externalization Full Steam Ahead - InferenceXnewsletter.semianalysis.com
- Real-time LLM Inference on Standard Datacenter GPUs (3,000 tokens/s per request)blog.kog.ai
- Redesigning the Inference Chip: From Nvidia GPU's Flaws to OpenAI Jalapeñozartbot.github.io
- The Inference Shift – Stratechery by Ben Thompsonstratechery.com
- AI Chip Architecturesjacobpeake.com
- Google "We Have No Moat, And Neither Does OpenAI"semianalysis.com
- Two different tricks for fast LLM inferenceseangoedecke.com
- My picture of the present in AI — LessWronglesswrong.com