Descriptions:
Fahd Mirza puts Qwen 3.8 Max through a dedicated multimodal evaluation, moving beyond text benchmarks to test what the model can actually do with images and video. The session opens with a genuinely difficult challenge: transcribing and comprehending a handwritten 1795 document with faded cursive ink and archaic spelling. Qwen 3.8 Max not only produced a clean line-for-line transcription including superscript abbreviations, but cross-referenced the named individual against historical records — going beyond pixel-reading into contextual reasoning.
Subsequent tests cover a numerical analysis convergence plot (where the model correctly identified fourth-order convergence slopes and explained the physical significance of floating-point roundoff dominating at low step sizes), a video comprehension challenge tracking three characters across cuts and inferring the filmmaker’s thematic intent, and a flag similarity test. Mirza notes an important architectural detail: Qwen 3.8 Max does not process video as a continuous stream — the platform samples frames and feeds them alongside audio, so the model reasons across a sequence of stills. Despite this, it successfully synthesized scene-level meaning rather than just describing individual frames.
The overall finding is that Qwen 3.8 Max’s vision performance is competitive with frontier multimodal models on the kinds of real-world tasks that trip up most systems: ambiguous handwriting, scientific charts, and implicit visual narratives. Mirza’s benchmark comparisons from Qwen’s own materials show the model at or near the top across visual understanding, document reading, and computer-use tasks.
📺 Source: Fahd Mirza · Published August 03, 2026
🏷️ Format: Benchmark Test







