DeepSeek V4-Flash Vision Is Out: Whale Opened Its Eyes

DeepSeek V4-Flash Vision Is Out: Whale Opened Its Eyes

More

Summary

Fahd Mirza puts DeepSeek V4 Flash Vision EXP through its paces in a hands-on review recorded on the night of its release. The experimental model adds image understanding to DeepSeek’s V4 Flash architecture, accepting photos, screenshots, charts, and documents alongside text input. Each image is capped at 384 tokens and billed at the same rate as V4 Flash text generation. Images can be submitted as base64, a URL, or via a new files API that allows uploading once and reusing by ID.

Mirza runs a structured set of tests: transcribing a dense mathematical equation into LaTeX (near-perfect, with one exponent placement error), parsing multilingual handwritten text across English, Urdu, Arabic, and Indonesian (partial โ€” Urdu missed entirely, Indonesian partially hallucinated), analyzing a three-series business chart with COVID-era supply chain data (strong reasoning, correctly identifying lead time as a leading indicator and citing exact figures like a 26-week peak and $620 billion drop), and extracting a four-column financial table from a low-quality image (approximately 97โ€“98% accurate, minor figure errors). A final roleplay test presents the model with a high-pressure fictional decision scenario, which it handles in-character with detailed prose.

The overall picture is a model with impressive chart reasoning and structured data extraction but meaningful gaps in multilingual OCR and fine-grained symbol transcription. Strong for business document analysis; less reliable for precise handwriting recognition across non-Latin scripts.


๐Ÿ“บ Source: Fahd Mirza ยท Published August 21, 2026
๐Ÿท๏ธ Format: Review

1 Item

Channels

1 Item

Companies