12:14 Tutorials1 month ago Ternary Bonsai 27B: Full Model, 2-Bit, Run on Phone or Laptop or GPU Fahd Mirza runs Bonsai 27B — a 27-billion parameter open-weight model from Prism ML based on Gemma 3 27B — on an NVIDIA RTX A6000 GPU... 0 comments 3.1K views
13:00 Coding & Dev Tools1 month ago Tess-4-27B + EAGLE-3: Local Reasoning, Nearly 2× Faster Fahd Mirza demonstrates running Tess-4-27B — a reasoning-capable fine-tune of Qwen 3.6 27B by Miguel de Icaza — locally on an NVIDIA... 0 comments 1.2K views
08:51 Benchmarks2 months ago Qwen3.6-27B with Thinking Cap on: Same Accuracy, 36% Less Thinking Fahd Mirza runs a live, unscripted comparison between Qwen3.6-27B and a community fine-tune called Thinking Cap — a variant of the sa... 0 comments 4K views
13:29 Coding & Dev Tools2 months ago Control What Your AI Agents Can Do: Archestra + Ollama Hands-On Fahd Mirza walks through Orchestra (stylized as Archestra), an open-source platform built for running AI agents safely in production,... 0 comments 1.3K views
08:09 Business & Strategy2 months ago Someone Fine-Tuned a Model on 10 Examples in 3 Minutes: Qwable-5 27B Coder Fahd Mirza uses a deliberately provocative model release — Qwable-5 27B Coder — as the centerpiece of a broader argument about credib... 0 comments 2.6K views
14:42 Tutorials2 months ago Qwen3.6 27B (Pi-Reasoning GGUF) – Fine-Tuned for Local Heavy AI Agent Fahd Mirza tests Pi-Reasoning, a community fine-tune of Qwen 3.6 27B built specifically for agentic coding — tasks like reading files... 0 comments 3.8K views
09:40 Benchmarks2 months ago DFlash Just Got Faster: 4x Speed with 160 tok/s Locally Fahd Mirza benchmarks DFlash with SGLang's new SpecV2 overlapping scheduler on an NVIDIA H100 80GB GPU, demonstrating a 4.3x throughp... 0 comments 2K views
13:26 Research & Benchmarks3 months ago Gemma4 12B vs Qwen3.6 27B — The Veteran vs The Newcomer Fahd Mirza runs a structured head-to-head comparison of Gemma 4 12B and Qwen 3.6 27B on the same Nvidia H100 80GB VRAM system, testin... 0 comments 4.3K views
32:57 Tutorials3 months ago Unsloth Studio is insane… fine-tune any AI model locally Unsloth Studio is a free, open-source desktop application that brings LLM fine-tuning to consumer hardware — and this video by David... 0 comments 8.3K views
10:06 Coding & Dev Tools3 months ago DFlash Leaves Qwen Territory – Gemma 4 31B Now Runs 5x Faster with Speculative Decoding Fahd Mirza demonstrates the first end-to-end deployment of Llama Box DFlash with Google's Gemma 4 31B model, following the merge of P... 0 comments 3.5K views