Summary
High Bandwidth Flash, or HBF, is a new memory approach aimed at the growing demands of serving large AI models, and it could challenge or complement today’s High Bandwidth Memory (HBM). Asianometry explains why memory makers are paying attention after the HBM boom.
The video starts with the fundamentals of inference: prefill, where a model processes the whole input in parallel and builds a KV cache, and decode, where tokens are produced one at a time and the workload becomes memory bound. Two trends are pushing memory hard. Model weights keep growing, with Kimi K3 at 2.8 trillion parameters compared with roughly a terabyte for K2.7 and 671 gigabytes for DeepSeek V3. Meanwhile, long-running agents create enormous KV caches that strain a context window of around one million tokens.
HBF stacks NAND flash dies connected with through-silicon vias. It offers 8 to 16 times the capacity of an HBM4 stack, about 512 gigabytes, with bandwidth of up to 3 terabytes per second. The tradeoffs are real: reads and writes are roughly two orders of magnitude slower than DRAM, NAND wears out after thousands of cycles, and power and heat are higher. The video also traces the idea’s history, from a 2019 High Bandwidth NAND paper to Sandisk’s 2025 HBF announcement and technical advisory board.
📺 Source: Asianometry · Published September 24, 2026
🏷️ Format: Deep Dive







