09:45 Tutorials5 months ago Llama.cpp Just Got MTP – Qwen3.6 27B Runs 2x Faster Locally with Two Flags Multi-token prediction (MTP) support has officially merged into the mainline llama.cpp repository—not a fork or custom branch, but th... 0 comments 3.1K views
16:50 Tutorials5 months ago Run HiDream-O1-Image Locally with ComfyUI Fahd Mirza walks through the complete local installation of HiDream-O1, a newly released 8-billion-parameter image generation model f... 0 comments 1K views
08:16 Tutorials5 months ago GroundedAI with Ollama – Universal Evaluation Interface for LLM Applications Fahd Mirza demonstrates how to install and run GroundedAI, an open-source evaluation framework built to detect hallucinations and mea... 0 comments 1K views
10:06 Coding & Builds5 months ago Toto 2.0: Datadog’s Observability AI Model – Full Install + Live Dashboard Datadog has released Toto 2, an open-source family of time series foundation models purpose-built for observability and DevOps teleme... 0 comments 1.1K views
10:54 Tutorials5 months ago Talkie: I Ran a 1930 AI Model Locally and Talked to People from the Past Fahd Mirza explores Talkie, a 13 billion parameter language model built by Alec Radford — the researcher behind GPT-2 — that was trai... 0 comments 624 views
09:22 Tutorials5 months ago DramaBox – Run Most Expressive TTS with Voice Cloning Locally Fahd Mirza takes a hands-on look at DramaBox, a newly released expressive text-to-speech model that can be run locally on consumer-gr... 0 comments 855 views
11:12 Benchmarks5 months ago Qwen3.6 27B Gets 20% Faster with MTP and llama.cpp Locally Fahd Mirza demonstrates how to enable multi-token prediction (MTP) on Qwen3.6 27B using ik_llama.cpp — a community fork of the popula... 0 comments 3.4K views
09:49 Tutorials5 months ago Building AI Evals for Real-World Problems Fahd Mirza walks through how to set up and run OpenAI's evals framework on a practical real-world task: classifying the causes of inv... 0 comments 434 views
11:00 Tutorials5 months ago NVIDIA Nemotron Elastic: 3-in-1 Elastic LLM Like Russian Dolls in One File NVIDIA's Nemotron Elastic model family packs three reasoning models — 30B, 23B, and 12B parameters — into a single checkpoint file us... 0 comments 1.5K views
09:15 Benchmarks5 months ago ZAYA1-VL-8B: Efficient Open Visual Intelligence – Run Locally Fahd Mirza puts ZAYA1-VL-8B — the new vision-language model from Zeffa — through its paces on an NVIDIA RTX 6000 with 48GB of VRAM, s... 0 comments 847 views