Summary
Fahd Mirza tests Meta’s newly released Muse Spark 1.3, a multimodal reasoning model targeting long-horizon agentic work with a 1 million token context window and pricing set at $1.25 per million input tokens and $4.25 per million output tokens. The headline news: Meta founder Mark Zuckerberg announced on X that an open-weight release is coming soon, a development Mirza highlights as significant for the broader AI ecosystem.
The primary test hands the model access to a live personal AWS account with a single goal: build an animated website from scratch, provision an S3 bucket and CloudFront distribution, and return a working public URL — all without step-by-step guidance. Mirza observes the model identifying independent tasks and executing them in parallel threads, creating the S3 bucket, drafting the HTML file, and configuring CloudFront simultaneously before combining the outputs. Despite heavy server throttling from Meta, the model completes the full end-to-end cloud deployment from a single text prompt.
Secondary tests cover vision and social reasoning — the model correctly reads a fragmented WhatsApp thread, identifies senders by bubble position, and catches the double meaning in an ambiguous message — along with a multi-stage chemistry and math buffer problem. On published benchmarks, Muse Spark 1.3 tops the field on professional tool use and agentic computer use, outperforming both GPT 5.6 and Claude Opus 5, with a particularly strong lead on million-token retrieval tasks where competitors score in the 60–70% range while Muse Spark approaches the high 90s.
📺 Source: Fahd Mirza · Published September 03, 2026
🏷️ Format: Review







