Summary
Louis-François Bouchard, Omar Solano, and Samridhi Vaid from Towards AI present an 80-minute conference workshop on context engineering in 2026, using their production AI tutor — built to support students in the Towards AI Academy — as the experimental testbed. The talk is grounded in real user interactions across a corpus exceeding 8 million tokens spanning course lessons, LangChain, LlamaIndex, OpenAI, and Claude Code documentation.
The team walks through their full retrieval pipeline: hybrid search combining a Cohere embedding model with BM25 keyword indexing to surface the top 30 candidate chunks, followed by reranking to the top 5, with a hard 100,000-token context cap. Configuration choices were validated through systematic recall experiments measuring whether the correct knowledge-base page was retrieved. Beyond retrieval, the workshop covers context compaction strategies, managing the stateless nature of LLMs across long tutoring sessions, and a novel ‘knowledge base browsing’ tool that lets the agent traverse the file system via bash commands when multi-document synthesis is needed.
All experiments and the AI tutor itself are open source, with a live HuggingFace Space linked in the slides. The session is practical throughout, offering directly applicable patterns for anyone building RAG-backed agents where context quality and cost are primary constraints.
📺 Source: AI Engineer · Published August 17, 2026
🏷️ Format: Deep Dive







