动察 Beating AI Flash: Nvidia CEO Jensen Huang disclosed this morning that GPT-6 Astra's training used approximately 100,000+ Grace Blackwell GPUs, forming a high-speed interconnected cluster via NVLink72. He also stated that the next batch of 400,000 GPUs is coming online, but did not specify the exact model or ownership.
Using Astra as a benchmark, China's leading AI model companies still show a significant gap in public computing power. ByteDance is currently the closest, having connected approximately 36,000 B200 chips via Malaysia this year. These chips belong to the same Blackwell generation as the GB200 used by Astra, but the computing power is deployed overseas. When ByteDance publicly introduced its domestic China clusters this year, it primarily relied on the China-specific H20 and H800, which are still based on the Hopper architecture.
Kimi was recently reported to have obtained approximately 20,000 Hopper GPUs through Alibaba. Bloomberg sources indicated they were specifically H200s, but Alibaba denied the H200 model designation. The H200 belongs to the previous Hopper generation; the B200 has already been upgraded to Blackwell. The GB200 further combines the Grace CPU with the Blackwell GPU, and NVL72 can place 72 GPUs within the same high-speed interconnect domain.
DeepSeek has not disclosed the full training hardware for V4. A leaked transcript from Liang Wenfeng's investor meeting stated that the company had approximately 20,000 H-equivalent computing power units in May, with most just arriving, and that future procurement would be "basically all Nvidia." Huawei's 950 production capacity allocated to DeepSeek is approximately 16,000 units. Liang Wenfeng stated that this batch of domestic chips is roughly equivalent to 4,000 Nvidia B-series units, only sufficient for training the current generation of models.
Meanwhile, the Blackwell used by Astra is no longer Nvidia's latest generation. Vera Rubin has already entered full-scale production. Nvidia estimates that Rubin can reduce the number of GPUs required to train large MoE models to one-quarter of what Blackwell requires.
In public records, no Chinese company has yet disclosed a case of using 100,000+ same-generation advanced GPUs for a single model training run. So the publicly visible gap between Chinese and American AI companies is no longer simply about the number of chips, but rather which generation of chips they can access, and how many can be concentrated on a single model at one time.

