Summary
Fahd Mirza puts Agents-A1 through its paces on a local Ubuntu system equipped with an NVIDIA RTX 6000 48GB GPU. Agents-A1 is a 35 billion parameter Mixture-of-Experts model released under Apache 2.0 by Intern Science, a Shanghai-based AI lab. The model is designed specifically for long-horizon agentic tasks — extended reasoning chains involving search, tool use, scientific analysis, and engineering — rather than maximizing raw parameter count.
What makes Agents-A1 architecturally interesting is its training pipeline, which Intern Science calls a Knowledge Action Infrastructure. Rather than recording only final answers, the system captures the entire chain of actions taken during problem-solving. Training proceeds in three stages: a broad fine-tune across domains, followed by deep specialist teacher models trained per domain (search, coding, science), and finally on-policy distillation back into the single 35B model. The result is a model that scales reasoning horizon rather than parameters.
Mirza tests the GGUF Q8 quantized version (consuming around 26GB VRAM) across two tasks: multimodal scientific reasoning using a wind tunnel aerodynamics diagram, and autonomous app generation from a screenshot using the Hermes agent framework. The model produces a correct aerodynamic explanation and a functional HTML application with responsive components. On official benchmarks, Agents-A1 leads its weight class on long-horizon search and scientific reasoning, though larger closed models like Gemini K2.6, DeepSeek V4 Pro, and GPT-5.5 still lead on raw coding tasks. The Q4_KM quantization is flagged as a viable option for systems with less VRAM.
📺 Source: Fahd Mirza · Published July 06, 2026
🏷️ Format: Review







