Robotics & Physical AIAug 27, 2026

Alibaba released Qwen3.8-Flash-Next, a sparse open-weight model that activates 6B parameters per token

Released on 26 August, Qwen3.8-Flash-Next is an open-weight multimodal mixture-of-experts model carrying a 125B backbone plus a 51B n-gram embedding layer (about 180B on the model card) and activating only 6B parameters per token. It ships a Gated DeltaNet and Qwen Sparse Attention hybrid, gated residuals and the Muon optimiser, with a 262,144-token native context extensible to a million. Alibaba frames it as a preview of the Qwen4 architecture and prices it at roughly a twelfth of its own flagship — listed on OpenRouter the same day at $0.15 per million input and $0.47 per million output tokens.

What it means A 6B active-parameter model with a million-token ceiling is the shape that makes long-context agent work affordable on hardware you already own, and at a twelfth of flagship pricing it resets what a routine call should cost.

Where it came from Qwen (Alibaba)

Back to the Stream