AMD Releases Instella-MoE-16B-A3B: A Fully Open Mixture-of-Experts LLM With 2.8B Active Parameters Trained On Instinct GPUs
MarkTechPost
Read Full Article at MarkTechPost →
Ad Slot — In-Article (728x90)
AMD released Instella-MoE-16B-A3B, a fully open Mixture-of-Experts language model trained from scratch on Instinct MI300X and MI325X GPUs. It holds 16B total parameters but activates only 2. 8B per token, using Gated MLA and FarSkip-Collective.
AMD published weights from every training stage, plus data mixtures, configs, and inference code. The post AMD Releases Instella-MoE-16B-A3B: A Fully Open Mixture-of-Experts LLM With 2. 8B Active Parameters Trained On Instinct GPUs appeared first on MarkTechPost.
This is a summary. For the full story, read the original article at MarkTechPost.
Original source: MarkTechPost