DeepSeek V4 is Here – Pro and Flash – Model That Made All GPU Clusters Obsolete

DeepSeek V4 is Here – Pro and Flash – Model That Made All GPU Clusters Obsolete

More

Summary

Fahd Mirza tests DeepSeek V4 on the day of its release, covering both the Flash variant (284 billion total parameters, 13 billion active) and the Pro variant (1.6 trillion total parameters). DeepSeek V4 introduces a fundamentally new architectural approach: a hybrid Compressed Sparse Attention (CSA) and Hierarchical Compressed Attention (HCA) system paired with a new optimizer and residual connection design. The result, according to DeepSeek, is million-token context processing at roughly 27% of the compute cost required by comparable models six months ago.

Mirza runs the model through a slime mold biological simulation โ€” a self-organized agent system that produces emergent network patterns โ€” a constrained multi-variable staff scheduling puzzle, and a multilingual translation task. The slime mold output renders with working decay rate controls, color themes, and real-time agent spawning that Mirza describes as genuinely lifelike. The scheduling puzzle correctly identifies an impossible constraint day and produces a clean roster for all solvable cases, handling deliberate ambiguities in the prompt without prompting for clarification.

The Flash variant already outperforms DeepSeek 3.2 across most benchmarks despite being significantly cheaper, while the Pro variant is positioned as the current strongest open-source model by multiple measures. For developers evaluating open-weight models for complex reasoning, long-context work, or self-hosted inference, DeepSeek V4’s architectural changes represent the most significant open-source release since DeepSeek R1.


๐Ÿ“บ Source: Fahd Mirza ยท Published April 24, 2026
๐Ÿท๏ธ Format: Review

1 Item

Channels

1 Item

Companies