‘Inference Speed Makes Markets Bigger,’ says Cerebras CEO

‘Inference Speed Makes Markets Bigger,’ says Cerebras CEO

More

Descriptions:

Cerebras Systems co-founder and CEO Andrew Feldman joins Bloomberg Tech to explain the newly announced AMD-Cerebras partnership around disaggregated inference infrastructure. The deal pairs AMD’s Helios server — optimized for prompt processing (the prefill phase, which is highly parallelizable and GPU-friendly) — with Cerebras’ wafer-scale engine for token generation, which Feldman describes as delivering the fastest inference throughput available. The combined Helios+Cerebras pipeline is slated to launch in the Cerebras cloud later in 2025, with broader availability to follow.

Feldman positions the architecture as a natural extension of Cerebras’ open-standards approach, noting that AMD has been an equity investor since mid and late funding rounds, and that the same open IO strategy has allowed partnerships with AWS Trainium as well. He draws a contrast with the NVIDIA-Groq integration (described as more tightly coupled) and argues that standards-based high-speed Ethernet interconnects let Cerebras plug into any major chip ecosystem.

The conversation also covers Cerebras’ recent IPO and how the company handled acquisition interest from major chip makers prior to its public listing. Feldman makes a broader market argument — that faster inference expands total spending rather than carving up a fixed market — and weighs in on the U.S.-China AI competition, noting Cerebras avoids HBM memory dependencies entirely, giving it a supply-chain and cost advantage as the industry debates dollar-per-token economics.


📺 Source: Bloomberg Tech · Published July 24, 2026
🏷️ Format: Interview

2 Items

Companies

1 Item

People