Descriptions:
Dan Biderman, co-founder and CEO of Engram — the AI memory startup that raised a $98 million seed round — joins the Latent Space podcast to explain why long context windows and compaction alone cannot solve the AI memory problem. The conversation, recorded in an unusual cooking-show format, covers the fundamental inefficiency of the KV cache: a 70B-parameter Llama model consumes roughly 80 GB of GPU HBM just to hold a single Wikipedia article in context, while the full model weights clock in at only ~140 GB. That mismatch, Biderman argues, makes current approaches both expensive and unreliable over long horizons.
Biderman draws on his background in computational neuroscience and prior work on LoRA at Mosaic ML (alongside co-founders from Stanford, Cornell, and Berkeley) to propose a complementary path: gradient-based neural memory traces that store information in weight space rather than token space. This is distinct from deterministic compaction — which evicts tokens with hard in/out decisions — and is closer to what some researchers call test-time training or test-time compute.
The near-term focus for Engram is token efficiency: getting models to reason over large contexts using fewer tokens and with less confusion deep into a session. Longer term, Biderman sees continual learning as the mechanism that enables truly long-horizon tasks in science, engineering, and defense — tasks that today remain out of reach because the computational and memory costs are simply too high.
📺 Source: Latent Space · Published July 13, 2026
🏷️ Format: Podcast







