Qwen3.8-Flash-Next releases open weights for a hybrid-attention architecture
Qwen/Qwen3.8-Flash-Next
AISummary
Qwen released open weights for Qwen3.8-Flash-Next, a 125B-parameter model with 6B activated, built on a new hybrid architecture with Gated DeltaNet and Qwen Sparse Attention. The model has a native 262,144-token context length, extensible to 1,000,000 tokens, and the source reports benchmark results across coding, agent, and vision tasks.
AIWhy it matters
The release pairs a new hybrid attention and gated residual architecture with open weights and benchmark results, giving architecture-focused readers a concrete case to compare against prior long-context designs.
Source: Qwen · new models on Hugging Face · huggingface.co