Summary
Zyphra, a San Francisco AI lab known for earlier releases like Zonos and ZR1, has returned with Zaya 1 (Zia) 8B — an open-source mixture-of-experts model featuring only 760 million active parameters out of 8.4 billion total. The model carries two notable distinctions: it was trained entirely on AMD hardware, reportedly the first model at this performance tier to do so, and its post-training pipeline uses a technique called Marovian RSA, a consensus-style ensemble method designed to boost reasoning accuracy on hard tasks.
Fahd Mirza installs Zia 8B locally using vLLM on an Nvidia RTX 6000 with 48GB of VRAM (consuming roughly 47GB loaded) and runs two demanding tests: a multi-step emergency aviation navigation problem requiring fuel burn, wind drift, and geometry calculations, and a real-time collaborative code editor request in Python with FastAPI. The model handles fuel math and wind vectors correctly but miscalculates a heading bearing — a mistake Mirza flags explicitly — illustrating both the model’s reasoning strengths and its limits.
Benchmark claims position Zia 8B as competitive with models many times its size, including Claude and Gemini 2.5 Pro on Hard difficulty evaluations, with performance approaching GPT-5 on select tasks. The video covers the CCA block attention mechanism, the MoE expert routing architecture, and the Marovian RSA boosting method in plain-language terms, making it a useful technical overview for developers evaluating efficient open-weight models for local or on-premise deployment.
📺 Source: Fahd Mirza · Published May 07, 2026
🏷️ Format: Review







