08:10 Coding & Builds3 months ago vLLM + PegaFlow: KV Cache That Survives Restarts (Hands-On) Fahd Mirza walks through a hands-on setup of PegaFlow, a new open-source tool from Novita AI that solves a persistent vLLM pain point... 0 comments 827 views
13:29 Coding & Builds3 months ago Control What Your AI Agents Can Do: Archestra + Ollama Hands-On Fahd Mirza walks through Orchestra (stylized as Archestra), an open-source platform built for running AI agents safely in production,... 0 comments 1.4K views
14:48 News & Opinion3 months ago Turbocharge Your Agent’s Retrieval with TurboQuant – Shashi Jagtap, Superagentic AI Shashi Jagtap, founder of SuperAgentic AI, presents at the AI Engineer conference on TurboQuant — a vector embedding compression algo... 0 comments 651 views
08:51 Tutorials3 months ago OpenJarvis + Ollama: Local AI Agent That Tracks Every Watt Fahd Mirza walks through the installation and hands-on testing of Open Jarvis, a newly released local-first personal AI framework dev... 0 comments 2.2K views
09:41 Coding & Builds4 months ago Qwen-AgentWorld: One AI Model That Simulates 7 Different Environments Fahd Mirza walks through a complete local installation and live demonstration of Qwen-AgentWorld, a novel "world model" from the Qwen... 0 comments 2.2K views
10:04 Coding & Builds4 months ago Ornith 1.0 9B: Self-Improving Model for Agentic Coding – Run Locally Fahd Mirza walks through a complete installation and evaluation of Ornith 1.0 9B, a newly released open-source model family built spe... 0 comments 3.7K views
09:00 Coding & Builds4 months ago SkillOpt: Microsoft’s New Way to ‘Train’ AI Agents: Run Locally Microsoft Research's SkillOpt takes a different approach to improving AI agent performance: instead of fine-tuning model weights, it... 0 comments 2.1K views
18:46 Deep Dives4 months ago You Might Not Need 50 Diffusion Steps — Ziv Ilan, Nvidia At the AI Engineer conference, Nvidia's Ziv Ilan — a researcher in Nvidia's AI labs team based in Paris — presents a practical framew... 0 comments 2.1K views
12:48 Tutorials4 months ago DiffusionGemma: 1100 Tokens/sec: Google’s Fastest Open Model Yet Locally Fahd Mirza installs and stress-tests Google DeepMind's DiffusionGemma — a 26-billion-parameter mixture-of-experts model that abandons... 0 comments 5.7K views
20:19 Coding & Builds4 months ago GPU Cloud Deployment Without Leaving Your IDE — Audry Hsu, RunPod Audrey Hsu, developer advocate at RunPod, demonstrates the company's new IDE-integrated GPU deployment tooling at AI Engineer, showin... 0 comments 213 views