Summary
Fahd Mirza independently tests Maple-Preview, a 20-billion-parameter sparse mixture-of-experts model from Deep Grove (an independent US lab) that uses ternary weights — a technique where each weight stores only one of three values (-1, 0, or +1) instead of a full 16-bit floating point number, shrinking the complete model to under 6 gigabytes. The experts and attention projection layers are ternary while embeddings and the output head retain higher precision, with 1.5 billion parameters active per forward pass across 256 experts with 8 firing per token.
The most notable finding is a near-complete absence of GPU benefit: running on an Nvidia H100 80GB, the model achieves 67 tokens per second on CPU-only mode and just 71 tokens per second with the full GPU — a gap of only 4 tokens per second. Deep Grove’s own published benchmark claims 218 tokens per second on a Mac Mini, meaning the Mac Mini outperforms an 80GB data center GPU by a factor of three. Mirza traces this to the fact that ternary weights require custom GPU kernels that currently exist only in Deep Grove’s own fork of llama.cpp and not in the mainline build, which defaults to standard floating-point paths.
The video also flags the absence of a technical report and model card, and notes that the Qwen 2 tokenizer reuse provides the only lineage hint available. Despite these red flags, Mirza concludes the model is a genuine technical experiment worth watching as ternary quantization tooling matures.
📺 Source: Fahd Mirza · Published August 09, 2026
🏷️ Format: Benchmark Test







