Beating AI News Flash: DeepSeek is once again scaling up its models. CEO Liang Wenfeng revealed to investors that the company is training a 2 trillion parameter model, and later plans to develop an 8 trillion parameter model. The current flagship V4-Pro has only 1.6 trillion parameters.
What does 8 trillion mean? Among the currently publicly available ultra-large models with open weights, Moonshot AI's Kimi K3 has reached 2.8 trillion parameters, already the world's largest open-source model. DeepSeek's planned 8 trillion parameter model would be nearly 3 times larger in scale.
However, more total parameters do not mean all parameters must be invoked in every run. Kimi K3 uses a MoE architecture, where only 104 billion of its 2.8 trillion total parameters are activated per token.

