19:19 Coding & Builds2 weeks ago How to Build Things with Jev & OpenJevs Sam Witteveen builds a working model router using Jev, a "system one" decision model from TypeSafe AI that returns typed, probability... 0 comments 3.5K views
14:30 Benchmarks4 weeks ago MiniCPM5-2B: The Best Sub-Agent Model Yet? Sam Witteveen examines MiniCPM 2B (technically 2.5B), the latest small model from OpenBMB, which claims to beat 4B-class models on fu... 0 comments 4.2K views
12:29 Reviews & Comparisons4 weeks ago Nex-N2.5 Mini: Multilingual, Multimodal, and Fully Agentic (Hands-On) Fahd Mirza provides a hands-on deployment walkthrough and capability test of Nex AGI's N2.5 Mini, the smaller model in the N2.5 agent... 0 comments 5.2K views
01:28:12 Deep Dives4 weeks ago Deep dive on LLM Inference at Scale — Harshul Jain, Audible & Tanmay Sah, Independent AI Researcher Harshul Jain, senior software engineer at Audible and author of an open-source LLM inference handbook, teams up with AI researcher Ta... 0 comments 1.2K views
13:50 Reviews & Comparisons1 month ago MiniCPM5 2B: SOTA or Scam? Let’s Test Locally Fahd Mirza puts MiniCPM5 2B — the latest small model from ModelBench and Tsinghua University's NLP lab — through a gauntlet of real-w... 0 comments 5K views
11:30 Reviews & Comparisons1 month ago Spark X2.5 4B: What a 4B Model Can and Can’t Do Locally Fahd Mirza puts the newly released Spark X2.5 4B through a full local deployment and capability evaluation on a 48GB VRAM GPU. The mo... 0 comments 1.9K views
12:54 Coding & Builds2 months ago Qwen3.8-27B in 2-Bit Quant: Escha-W2 Build Locally without Loss Fahd Mirza demonstrates Escha-W2, Asha Labs' aggressively quantized version of Qwen3.8 27B, which compresses the model from 50GB at f... 0 comments 4.2K views
09:20 Benchmarks2 months ago DFlash 2: Qwen3.8-27B at 2× Speed – Live Benchmark Locally Fahd Mirza benchmarks DFlash 2, a community-developed speculative decoding enhancement for the Qwen3.8-27B model, running live on an... 0 comments 6.5K views
08:39 News & Opinion2 months ago Why You Shouldn’t Run Qwen3.8 2.4T Locally Fahd Mirza delivers a frank assessment of Qwen's newly released Qwen3 2.4-trillion-parameter open-weight model, making the case that... 0 comments 1.3K views
19:50 News & Opinion2 months ago Taking Reinforcement Learning Cross Datacenter — Nan Jiang, Modal At an AI Engineer conference, Nan Jiang from Modal presents a technical architecture for running reinforcement learning post-training... 0 comments 300 views