header-langage
简体中文
繁體中文
English
Tiếng Việt
한국어
日本語
ภาษาไทย
Türkçe
Scan to Download the APP

DeepSeek plans to train an 8-trillion-parameter model, nearly triple the size of the largest existing open-source models.

Beating AI News Flash: DeepSeek is once again scaling up its models. CEO Liang Wenfeng revealed to investors that the company is training a 2 trillion parameter model, and later plans to develop an 8 trillion parameter model. The current flagship V4-Pro has only 1.6 trillion parameters.


What does 8 trillion mean? Among the currently publicly available ultra-large models with open weights, Moonshot AI's Kimi K3 has reached 2.8 trillion parameters, already the world's largest open-source model. DeepSeek's planned 8 trillion parameter model would be nearly 3 times larger in scale.


However, more total parameters do not mean all parameters must be invoked in every run. Kimi K3 uses a MoE architecture, where only 104 billion of its 2.8 trillion total parameters are activated per token.

举报 Correction/Report
Correction/Report
Submit
Add Library
Visible to myself only
Public
Save
Choose Library
Add Library
Cancel
Finish