The Inference Frontier: from 100 to 10,000 tokens per second — Sean Lie, Cerebras CTO

The Inference Frontier: from 100 to 10,000 tokens per second — Sean Lie, Cerebras CTO

More

Summary

Sean Lie, CTO of Cerebras Systems, joins the Latent Space podcast the day after Hot Chips to unpack the CS4 and forthcoming CS5 wafer-scale inference chips — and to contextualize OpenAI’s Jalapeno announcement within the broader inference hardware race. Lie describes Hot Chips 2026 as evidence of the golden age of semiconductor innovation, with simultaneous progress across chip design, interconnects, system architecture, optics, and AI-assisted tooling that has no historical precedent.

CS4 delivers twice the power and interconnect bandwidth of the previous Cerebras generation at half the latency, with the company claiming the top position in ultra-fast inference. CS5, expected next year, targets up to 10,000 tokens per second — a regime GPU architectures cannot reach and one Lie argues will enable qualitatively new product categories. He frames Jalapeno and CS5 as complementary rather than competing: Jalapeno is optimized for throughput, CS5 for latency, and OpenAI — Cerebras’s largest customer — is positioned to run both as a combined fast-inference portfolio. Lie also discusses how Cerebras is using OpenAI’s own AI tooling to accelerate its chip design and software development, creating a collaborative loop between the two companies.

The conversation covers Groq’s SRAM-based chip announcement, the performance-per-watt debate dominating offstage discussions at the conference, prefill/decode disaggregation opportunities, and why Nvidia’s CUDA software moat makes it difficult to displace even when a challenger chip outperforms on raw specs. Essential listening for anyone tracking the inference hardware landscape heading into 2027.


📺 Source: Latent Space · Published September 02, 2026
🏷️ Format: Interview

1 Item

Channels

2 Items

Companies