Summary
Fahd Mirza reviews and live-tests Ling-3.0-flash-VL, a new open-weights multimodal vision model from Inclusion AI released under an MIT license and currently free on their API. The model features 124 billion total parameters but activates only 5.5 billion per token via a mixture-of-experts architecture, making it significantly more efficient than its parameter count suggests. Notably, it was post-trained on Kimi K3, Moonshot AI’s frontier model, and the influence shows — Mirza repeatedly describes the reasoning style as having a distinct “Kimi flavor.”
Architecturally, Ling-3.0-flash-VL routes every image or video through a vision encoder before projecting it into the same embedding space as text, then passes it through 42 transformer layers arranged in a 5-to-1 pattern: five Kimi Delta Attention layers for efficient deep reasoning followed by one Multi-head Latent Attention layer for long-range context. A video-aware positional encoding (VideoRoPE) tracks spatial and temporal position simultaneously, giving the model genuine motion understanding rather than frame-by-frame description.
Benchmark results show a jump from 38 to 42 on a multimodal index compared to the text-only Ling-3.0-flash — suggesting vision integration improved general reasoning, not just added a new modality. Tests in the video include generating an animated 3D HTML simulation of a döner kebab from a photo, decoding workplace humor with double meanings in a WhatsApp screenshot, and translating a single English phrase into 80 languages. Performance is strong on major languages (Mandarin, Hindi, Arabic, Japanese, Korean) but degrades noticeably on lower-resource African and Central Asian languages.
📺 Source: Fahd Mirza · Published September 10, 2026
🏷️ Format: Review







