Multi-GPU Kernels, Intelligence per Watt, Heterogeneous Inference, and More | YC Paper Club

Multi-GPU Kernels, Intelligence per Watt, Heterogeneous Inference, and More | YC Paper Club

More

Descriptions:

This YC Paper Club special edition gathers Stanford researchers and senior engineers for a dense technical panel on GPU kernel optimization, chip specialization, and the evolving architecture of AI inference infrastructure. Hosted at Y Combinator Mountain View, the session features Stuart Soul (research scientist at Cursor, trains Composer models), John (Stanford CS PhD focused on intelligence-per-watt and intelligence-per-joule efficiency metrics), Mark (former PyTorch maintainer, co-founder of GPU Mode, now co-founding Core Automation with OpenAI VP of Research Jerry Tworek), Misha (AI infrastructure veteran from NVIDIA and Meta), and Brennan (RL researcher specializing in GPU-accelerated simulation).

The panel’s central argument is that specialization is now the defining force at every layer of the stack. Speakers explain why training and inference data centers have fundamentally divergent hardware requirements, how Google’s TPU v8 Zebrafish and Sunfish variants signal the beginning of deep ASIC divergence, and why batch-size-one inference for latency-critical applications like voice agents demands purpose-built silicon. The current hybrid approach — NVIDIA for prefill, Cerebras for decode — is framed as an early sign of what will become a much more fragmented chip landscape.

Mark’s segment dives into AI-written GPU kernels, drawing on the KernelBench competitive benchmarking platform and KernelGuard (an anti-cheating counterpart) to assess how well LLMs currently generate production-quality CUDA code. The session also covers remaining algorithmic headroom on the software side — including model routing strategies that avoid invoking full large models for trivial queries — making it essential viewing for engineers working on inference systems, custom hardware, or kernel-level optimization.


📺 Source: Y Combinator · Published July 29, 2026
🏷️ Format: Keynote Launch

1 Item

Channels