I Tested Ornith’s New 35b MoE on Mac — Here’s How It Goes

I Tested Ornith’s New 35b MoE on Mac — Here’s How It Goes

More

Summary

Bart Slodyczka runs Ornith 1.5’s 35B mixture-of-experts model through a series of practical one-shot tests on an Apple M3 Ultra Mac Studio. Ornith 1.5 launched in three sizes — a 9B dense model, the 35B MoE tested here, and a 397B flagship — and Slodyczka uses a 4-bit OQ4E quantization optimized for the MLX engine. The model loads at 21.6 GB, with an additional 21.5 GB required for a full 260,000-token context window, totaling 43 GB on device.

Performance benchmarks show 120 tokens per second at zero context with multi-token prediction (MTP) enabled, dropping to 76 tokens/sec at 20,000 tokens — significantly faster than the 25 tokens/sec Slodyczka recorded for Qwen 3.8B at 70,000 tokens. An 8-bit variant with MTP clocks in at 107 tokens/sec, and the baseline 4-bit without MTP lands at 83 tokens/sec. Vision tests involve extracting structured data from generated invoice images of varying density, while code generation tests cover a tower defense game, an Excel-style spreadsheet calculator with formulas, and a first-person Three.js game.

A notable methodology finding: adding a plan-then-verify step to the prompt — asking the model to outline, build, then self-check — helped the 4-bit model catch minor syntax errors (misplaced semicolons, bracket mismatches) within a single generation, improving one-shot success rates. Full prompts, results, and model outputs are published to Slodyczka’s GitHub for independent replication. He rates Ornith 1.5 as a clear improvement over Ornith 1.0 for local inference on Apple Silicon.


📺 Source: Bart Slodyczka · Published August 20, 2026
🏷️ Format: Benchmark Test

1 Item

Channels

1 Item

People