Alibaba’s Qwen Team Releases Qwen3.8-Flash-Next: A 125B Multimodal MoE With 6B Active Parameters Previewing the Qwen4 Architecture
MarkTechPost
Read Full Article at MarkTechPost →
Ad Slot — In-Article (728x90)
We look at Qwen3. 8-Flash-Next, Alibaba's open-weight multimodal Mixture-of-Experts model and an early preview of the Qwen4 architecture.
We break down where the 180B parameters actually sit: a 125B backbone, a 51B N-gram embedding table, and a 4B multi-token prediction module, with only 6B active per token.
This is a summary. For the full story, read the original article at MarkTechPost.
Original source: MarkTechPost