❯ Alibaba Open-Sources Qwen3.8-Flash-Next — 125B Model Activates Only 6B Parameters per Token
RELEASEAlibaba’s Tongyi Qianwen team open-sourced Qwen3.8-Flash-Next on August 26, positioning it as a preview of the next-generation Qwen4 architecture. The parameter mix is unusual: beyond the 125B backbone, it carries a 51B-parameter N-gram embedding table and a 4B-parameter multi-token prediction module, yet each token actually activates only 6B parameters. Bloomberg reported that the production version that followed, Qwen3.8-Flash, targets parity with Claude Opus 4.6 and DeepSeek V4-Flash.
SAVINGSThe official line is that training cost is roughly one-ninth that of the previous Qwen3.7-Plus, while coding and office-task capabilities have pulled ahead. Four changes are stacked together: three of every four layers use a gated incremental network, with the fourth using sparse attention at micro-block granularity; 20 million bigram and trigram phrases serve as a lookup-table embedding, which can also be offloaded to host memory on Nvidia devices; the residual stream is widened with gating; and Muon and AdamW are paired as dual optimizers. Native context is 262K tokens, extendable to 1 million.
PITFALLSThe license is qwen-community-1.0, not Apache 2.0 — read the terms before commercial use. The weights are not light either: the FP8 checkpoint is 172.78 GiB, and community testing shows it needs multiple GPUs rather than a single workstation. Production API pricing is $0.16 per million input tokens and $0.47 per million output tokens. In the same week, Zhipu benchmarked against Opus 4.8 and Alibaba against Opus 4.6, with both switching their reference point to the same closed-source rival. For overseas developers, the selection question is no longer whether the model can run, but how much the per-million-token bill differs.
▪ SIGNALThe one-ninth training-cost figure says more than any benchmark score about which battle the next-generation Qwen aims to win.