Cancel your subscriptions, Ox-Alpha is here! (GLM 5.3 Flash)

Cancel your subscriptions, Ox-Alpha is here! (GLM 5.3 Flash)

More

Summary

Matthew Berman covers GLM 5.3 Flash, the latest open-weights model from Chinese AI lab ZAI that briefly appeared on OpenRouter under the mysterious alias “Ox Alpha” before its identity was revealed. The model is a 320-billion-parameter mixture-of-experts architecture with only 18 billion parameters active at any time, making it fast and cheap to run while delivering benchmark results competitive with much larger closed models.

Berman walks through performance numbers across Terminal Bench, Deep Suite, and GDP-Val, where GLM 5.3 Flash places near Claude Opus 4.8 on coding and agentic tasks and tops the OpenAI real-world knowledge benchmark by a wide margin. Pricing comes in at roughly one-tenth the cost of the previous GLM 5.2, putting it in direct competition with GPT 5.6 Luna on the cost-efficiency curve. One caveat: the model uses significantly more output tokens per task than Luna, which affects total cost calculations for high-volume workloads.

The most striking detail in the video is a report from SemiAnalysis indicating ZAI served 100 trillion tokens per day entirely on Chinese-made chips — no Nvidia hardware involved. Berman frames this as a signal that China is co-designing models and silicon together, creating a vertically integrated AI stack that could sustain near-frontier intelligence at a fraction of the cost. The video includes live demos and a sponsor segment for Higsfield’s unified AI image/video API.


📺 Source: Matthew Berman · Published August 29, 2026
🏷️ Format: Review

1 Item

Channels

1 Item

Companies