Summary
Why are inference platforms that mostly serve open-weight models earning some of the largest revenue run rates in AI? This bycloud deep dive examines the economics behind inference-as-a-service, noting that five of the top ten AI startups with at least $500 million in run rate, excluding the big three labs, are inference providers.
The video sorts the market into four groups: hyperscalers such as Amazon, Microsoft, Google and Oracle; neoclouds like Nebius, CoreWeave and Crusoe; full-stack inference platforms including Fireworks, Together AI, Baseten and DeepInfra; and custom-silicon players such as Groq, Cerebras, SambaNova and Etched. It then explains what inference actually involves, from prefill and decoding to the role of the KV cache.
Key optimizations are covered in depth. Prompt caching can make cache-hit tokens dramatically cheaper, and DeepSeek V4 Pro pricing illustrates a gap of about 120 times between hit and miss tokens. The video also unpacks how DeepSeek’s reported 73.7 thousand tokens per second figure depended on more than half of input tokens being cache hits, and explains how batching raises GPU utilization. Viewers come away understanding how these companies make money and where the sector may be heading.
📺 Source: bycloud · Published September 29, 2026
🏷️ Format: Deep Dive







