Deepseek did it again…

Deepseek did it again…

More

Summary

DeepSeek has released V4.1 Flash, a 552-billion-parameter mixture-of-experts model that the company benchmarks against Anthropic’s Claude Opus 5 and OpenAI’s GPT 5.6 Soul. The model activates only 8 billion parameters for input and 16 billion for output, with a KV cache footprint 13 times smaller than DeepSeek V4 Pro and 4x smaller than the prior V4 Flash. On Terminal Bench 3.0 it scores 30 — second only to Opus 5 — and on Deep Suite achieves 74.2, reportedly beating both Opus 5 and GPT 5.6 Soul.

Pricing is aggressively tiered: 15 cents per million input tokens during off-peak hours ($0.30 peak) and 60 cents per million output tokens off-peak ($1.20 peak), with cache-hit pricing dropping to fractions of a penny. The model is fully open weights — downloadable and self-hostable without routing data through DeepSeek’s servers — with quantized versions expected to eventually run on consumer hardware.

Matthew Berman covers the release with measured skepticism: his own hands-on tests show the model underperforming its published benchmarks, a recurring gap with efficiency-focused open-source releases. He frames V4.1 Flash within a broader pattern where open-source Chinese models trail the absolute frontier by roughly six months in capability but lead dramatically on cost-efficiency, creating a bifurcated market where penny-per-token inference handles commodity tasks while $50-per-million-token frontier models handle the cases where accuracy is non-negotiable.


📺 Source: Matthew Berman · Published September 11, 2026
🏷️ Format: News Analysis

1 Item

Channels

1 Item

Companies