Summary
A mysterious model called 0x Alpha surfaced on OpenRouter with an extraordinary set of specs: a 1-million-token context window, multimodal input (text, image, and video), zero data retention, and a claimed throughput of 100 trillion tokens per day — all offered for free. TheAIGRID breaks down the community’s rapid-fire investigation into what this model actually is and what its sudden appearance means for the AI landscape.
The most credible theory, supported by jailbreaker ‘Ply the Liberator’ and prediction markets sitting at 90% confidence, is that 0x Alpha is GLM 5.3 Flash from Z.AI — the Chinese lab behind the ChatGLM family. Benchmark results on the DBSE place it roughly on par with GPT-5.6 soul-mid, though results vary across testers and private evaluations show it underperforms on harder tasks. If the GLM identification holds, that means a model competitive with mid-tier GPT-5.6 could soon be self-hosted on hardware like NVIDIA’s DGX Spark.
Beyond model identity, the video raises a deeper question: who is providing the compute for what appears to be unlimited free inference? Theories include a new undisclosed US investment partner or a strategy to harvest vast quantities of coding requests to fuel continuous learning. The combination of frontier-grade capability, multimodal support, and near-zero cost — if sustained — would represent a significant shift in the economics of AI access.
📺 Source: TheAIGRID · Published August 25, 2026
🏷️ Format: News Analysis







