Descriptions:
Matthew Berman covers the release of Grok 4.6, xAI’s latest model and a significant step up from its predecessor Grok 4.5. The video walks through the full benchmark picture: Grok 4.6 claims the top spot on GDP-val (OpenAI’s agentic knowledge-work benchmark), lands in a competitive third place on DeepSweet at 65.9 versus GPT 5.6 Soul’s 73, jumps from 15% to 26% on TerminalBench, and dominates Harvey Lab’s legal benchmarks at 15.8% compared to Soul’s 2.5%. On the Artificial Analysis Intelligence Index, the model reaches an overall score of 61, tying with GPT 5.6 Soul and sitting just behind Fable 5.
Berman emphasizes that raw benchmark scores only tell half the story, spending considerable time on cost-per-task analysis using Artificial Analysis data. Grok 4.6 is priced at $2 per million input tokens and $6 per million output tokens โ a fraction of GPT 5.6 Soul’s cost โ and the cost-per-task chart shows meaningful efficiency gains over Grok 4.5 despite a roughly doubling of compute spend.
The broader narrative the video builds is strategic: SpaceX’s acquisition of Cursor has given xAI the coding-focused flywheel that previously propelled Anthropic and then OpenAI to dominance among developers. Berman argues Grok 4.6 marks the moment xAI should be considered a legitimate third major U.S. AI lab. The model is available today via Cursor, the Grok Build platform, the xAI API, OpenRouter, Vercel, and Cloudflare.
๐บ Source: Matthew Berman ยท Published August 13, 2026
๐ท๏ธ Format: Review





