Perceive Beating AI News Flash: Alibaba Thousand Whys releases Qwen3.8-Flash and simultaneously open-sources the model weights Qwen3.8-Flash-Next. The open-source version added a "Next," indicating that it is the first to use the next-generation Qwen4 architecture.
This is a multimodal MoE model, with the main model having 125B parameters and only 6B activation parameters. Additionally, it has added 51B N-gram Embedding, equivalent to a large "look-up memory," trading more storage for less computation.
In official evaluations, SWE-bench Pro outperforms Claude Opus4.6 by 9.1 points, JobBench by nearly 20 points; AndroidWorld by 22.5 points, and MathVision by 25.1 points. The model natively supports 262K contexts, extendable to 1M.
The price is also kept very low. The production version is simultaneously launched on Qwen Cloud, supporting 1M contexts and tool calls by default. It charges 1 euro for every million Token input and yields 3 euros in output.
Officially stated, Qwen3.8-Flash achieves a performance close to Qwen3.7-Plus but with a training cost of only about 1/9. Thousand Whys aims to use less computation to make the capabilities close to the flagship model more affordable.

