Summary
OpenAI has unveiled its first custom AI chip, codenamed Jalapeno, and independent semiconductor research firm SemiAnalysis visited the lab and ran their own tests — finding that it outperforms Nvidia’s Blackwell and Google’s TPUs on performance per watt across nearly all scenarios. At concurrency-one on DeepSeek R1, Jalapeno hits over 700 tokens per second per user, and the chip moved from initial design to tape-out in just nine months, with AI playing a direct role in the design process.
The chip is an ASIC — application-specific integrated circuit — built exclusively for large language model inference, not a general-purpose GPU. What makes the story particularly striking is how OpenAI sidestepped the software challenge that has stymied every other Nvidia competitor: rather than building a developer-friendly ecosystem to rival CUDA’s two decades of tooling, OpenAI writes Jalapeno’s kernels in assembly-level code using an internal version of Codex. Each kernel runs to thousands of lines, backed by correctness checks and a custom sanitizer, with the work shifting from human-in-the-loop to fully automated.
Wes Roth walks through the SemiAnalysis report, explains what kernels are and why the CUDA moat has been so hard for AMD and Google to overcome, and explores what it means that OpenAI chose to make the chip AI-programmable rather than human-friendly. The broader implication: if AI can design chips and write the low-level code to run them, the competitive dynamics of AI infrastructure may shift faster than the industry anticipated.
📺 Source: Wes Roth · Published August 26, 2026
🏷️ Format: News Analysis







