Ornith 1.5 35B-A3B: Self-Improving Model Goes General-Purpose: Run Locally

Ornith 1.5 35B-A3B: Self-Improving Model Goes General-Purpose: Run Locally

More

Descriptions:

Fahd Mirza walks through deploying Ornith 1.5 35B-A3B — the latest release from Deep Reinforce — on a single Nvidia A100 80GB GPU using vLLM. The model is a mixture-of-experts architecture that activates only 3 billion parameters per token, yet outperforms Qwen 3 35B across coding and agentic benchmarks and closes the gap on Qwen’s much larger 397B model. The video covers VRAM consumption (approximately 74GB), KV cache configuration, and the practicalities of serving the model locally.

The centerpiece demo gives the model an open-ended agentic task: a working full-stack crypto tracker application (FastAPI backend, WebSocket price feed, Redis history, Docker Compose) with no explicit instructions — just a goal to build a real-time price alert feature end-to-end and self-verify it works. The video shows Ornith 1.5 autonomously exploring the codebase, planning, writing code across multiple files, detecting port conflicts, running its own tests, and iterating until satisfied.

Mirza also explains the model’s GRPO-based self-improvement training loop — scaffold construction, multi-rollout generation, reward feedback — and how the live demo mirrors that same generate-plan-execute cycle. A second test demonstrates targeted tool use including automated pass/fail test reporting. Viewers interested in running capable open-weight agentic models on a single high-VRAM GPU will find this a practical end-to-end walkthrough.


📺 Source: Fahd Mirza · Published August 19, 2026
🏷️ Format: Hands On Build

1 Item

Channels