Descriptions:
Fahd Mirza takes a blind capability look at “Aux Alpha,” a mystery model quietly available on several AI providers under a code name, with no official lab, no announcement, and no branding. The model is listed as supporting 1 million tokens of context, multimodal inputs (text, images, and video), and is built for reasoning, coding, and long agentic work — currently free, with the trade-off that user prompts are used for the provider’s post-training pipeline.
To test real-world capability, Mirza deploys the model via Hermes Agent against Silo Trace — a live five-service Docker application simulating a feed mill monitoring dashboard with a deliberately introduced quality-check bug that inverts pass/fail logic for batch specifications. Without being told where the bug is or how many exist, Aux Alpha locates and fixes the flipped comparison in roughly 17 minutes, spending the majority of that time on testing and verification. Mirza notes it used fewer tokens and less wall-clock time than previous frontier models he has benchmarked on similar tasks.
Additional tests probe multimodal reasoning through an image-based high-stakes dilemma prompt and an 80-language evaluation requiring country identification, famous person selection, and authentic native-language quotes from each. Results across all three tests suggest a capable instruction-following and long-context reasoning model. Users are explicitly cautioned against sharing sensitive data given the provider’s stated retraining data collection policy, and production deployment is not recommended until provenance is established.
📺 Source: Fahd Mirza · Published August 22, 2026
🏷️ Format: Review






