AMD Releases First Ever AI model: Instella-MoE-16B-A3B-Think

AMD Releases First Ever AI model: Instella-MoE-16B-A3B-Think

More

Descriptions:

AMD has released Instella-MoE-16B-A3B-Think, its first-ever mixture-of-experts language model, trained entirely on AMD Instinct MI300X and MI325X hardware using AMD’s own open-source PRISM and MILES frameworks. Fahd Mirza focuses the video not on benchmark bragging but on two specific architectural contributions AMD has introduced.

The first, an attention quality gate called MLA, adds a learned scaling mechanism that sits after the attention operation and suppresses low-signal context before it propagates forward β€” effectively teaching the model to ignore irrelevant parts of its context window. The second, FarSkip Collective, is a distributed training optimization that rearranges GPU communication so data transfers happen in the background while computation continues, reducing idle time. AMD reports FarSkip cut pre-training time by 12.7% and reduced time-to-first-token at inference by up to 39.2%, though Mirza is careful to flag these as vendor-published, unverified figures.

Mirza also offers candid criticism of AMD’s benchmark comparisons, noting the model is measured mostly against older or lesser-known baselines like OLMo and GLM 3.5 rather than current frontier models, which makes the results harder to interpret. He spent three to four hours unsuccessfully attempting to run the model on an Nvidia GPU before giving up. Despite the lukewarm benchmark story, he sees FarSkip as a potentially influential idea that future models β€” especially code-focused ones β€” are likely to adopt.


πŸ“Ί Source: Fahd Mirza Β· Published July 28, 2026
🏷️ Format: Deep Dive

1 Item

Channels

1 Item

Companies