Memory Harnesses for Long-Running Research Agents — Stefania Druga, Sakana.ai

Memory Harnesses for Long-Running Research Agents — Stefania Druga, Sakana.ai

More

Summary

Stefania Druga, research scientist at Sakana AI in Tokyo, presents her work on memory harnesses designed to prevent context rot in long-running research agents. Speaking at the AI Engineer conference, she frames memory not as a passive database but as a write-manage-read control loop that wraps around the model itself — a design distinction she argues is critical as the industry moves toward longer-horizon agentic tasks.

Druga runs all her experiments locally on an M3 Ultra Mac with 96 GB RAM, using two quantized models: Qwen 27B at 4-bit and DeepSeek V4 Flash. Her harness evaluates four recall modes — no memory, vector RAG, a ranked decisions ledger, and an oracle baseline — across tasks from a literature review corpus and the X-Bench long-horizon benchmark. Across 68 benchmark questions, she finds the ranked-recall ledger outperforms vanilla RAG and even outperforms a gated memory strategy that asks the model to decide when it needs memory.

One counterintuitive finding stands out: the oracle condition, which feeds the model the provably correct memory entry, still fails to hit maximum accuracy because the model can receive correct context and still ignore or misinterpret it. This separates retrieval quality from utilization quality as two distinct problems. The talk closes by noting that for tasks that fit entirely within context, a memory harness adds cost without capability gains — making harness design a question of task scope as much as architecture.


📺 Source: AI Engineer · Published August 12, 2026
🏷️ Format: Deep Dive

1 Item

Channels

1 Item

Companies