header-langage
简体中文
繁體中文
English
Tiếng Việt
한국어
日本語
ภาษาไทย
Türkçe
Scan to Download the APP

NVIDIA: Qwen 3.8 Flash-Next Achieves Over 16,000 Tokens per Second per Card on GB 300 NVLink v7.2

Dynamic Beating AI News: NVIDIA has released an article stating that Alibaba's latest preview model, Qwen3.8-Flash-Next, has received support from the NVIDIA GB300 NVL72 platform. This model has a total parameter size of 176 billion, with each Token activating around 6 billion parameters. It natively supports contexts of 262,000 Tokens and can be extended to 1 million Tokens through YaRN, targeting applications such as intelligent programming, document processing, and tool invocation for long-context Agents.


NVIDIA has indicated that Qwen3.8-Flash-Next adopts a hybrid architecture of Gated DeltaNet (GDN) and Qwen Sparse Attention (QSA) to reduce computation and KV cache overhead in long-context scenarios. Testing has shown that on the GB300 NVL72, this model achieves a single GPU throughput of over 16,000 Tokens/second and a single-user throughput of over 200 Tokens/second. It also supports inference frameworks such as SGLang, vLLM, and TensorRT-LLM.

举报 Correction/Report
Correction/Report
Submit
Add Library
Visible to myself only
Public
Save
Choose Library
Add Library
Cancel
Finish