Descriptions:
Fahd Mirza puts Ling-3.0-Tiny — a new small model from Chinese lab Inclusion AI — through a live agentic evaluation rather than relying solely on its published benchmark claims. The model is a mixture-of-experts architecture with 7.9 billion total parameters but only 1.3 billion active per token, supports 262,000-token context, and is currently available free via API with open weights coming soon.
For the first test, Mirza feeds it a broken full-stack iron ore price validation app (FastAPI backend, SQLite, plain HTML/JS frontend) with deliberately planted bugs and asks the Hermes agent framework to fix everything from a single prompt. Ling-3.0-Tiny identifies and corrects all planted bugs across both front and back ends with no hand-holding, and the repaired app validates correctly in the browser — a strong result for a sub-2B active-parameter model. A follow-up hallucination test asks it about five nonexistent entities; the model declines to fabricate answers, consistent with its claimed benchmark lead on honesty metrics.
Mirza contextualizes Ling-3.0-Tiny within a competitive gap: Alibaba’s Qwen team has largely pulled back from the small-model segment, and Inclusion AI — previously known for flagship trillion-parameter models — appears to be deliberately targeting that opening. Benchmark comparisons show the model leading Qwen 3.5 at 4B and 9B, and both Gemma 4 sizes, on agentic and banking agent tasks, while conceding ground on broad knowledge and long-context retrieval.
📺 Source: Fahd Mirza · Published August 06, 2026
🏷️ Format: Review






