Summary
Fahd Mirza demonstrates MemPalace, a free open-source local memory system for AI that scored 96.6% on the LongMemEval benchmark — reportedly the highest published score on that evaluation — without making any API calls or relying on cloud infrastructure. The entire setup runs locally using Ollama serving Gemma 4 31B on an NVIDIA RTX 6000 with 48GB VRAM, keeping all data private and offline.
MemPalace organizes stored information using a spatial metaphor: wings (top-level containers for people or projects), rooms (specific topics like authentication or deployment), halls (connections between rooms within the same wing), and tunnels (automatic cross-wing links when the same topic appears in multiple projects). Each room maintains a closet (summary) alongside drawers storing original verbatim content — the design choice Mirza credits for the high recall score, since no information is lost through aggressive summarization. In the demo, MemPalace is initialized against a production-grade FastAPI template by Sebastián Ramírez, indexing 181 files into ChromaDB as over 1,000 drawers in seconds.
The integration with Ollama involves a simple script that first queries MemPalace for relevant context, then passes the results to Gemma 4 as grounded input. A test query about authentication architecture instantly surfaces JWT token creation logic, Argon2 password hashing, timing attack prevention, and the bcrypt-to-Argon2 migration path from a single natural language prompt. For developers looking to add persistent, semantically searchable memory to local AI workflows without external dependencies, this is a practical and well-structured introduction.
📺 Source: Fahd Mirza · Published April 16, 2026
🏷️ Format: Tutorial Demo







