Qwen/Qwen3.8-Flash-Next is an experimental open-weight model previewing the architecture planned for Qwen4. The release pairs Gated DeltaNet with Qwen Sparse Attention (QSA), which operates at the micro-block level to reduce long-context latency, and introduces a Gated Residual design. Weights are compatible with Transformers, vLLM, SGLang, and Docker Model Runner. Qwen Cloud offers managed inference, with Qwen3.8-Flash as the production version featuring 1M context length by default.
No score is assigned. Sources and their independence are shown in the citation chain below.