BlockBeats News, July 19th - Semiconductor and AI research firm SemiAnalysis published an article stating that although about three-quarters of Kimi K3's network layer adopts KDA, which can reduce KV cache transfer bandwidth by up to 10 times compared to a full-scale global attention model, this does not mean that the AI network switch market will shrink significantly.
Kimi K3 has 2.8 trillion parameters, and even with MXFP4, each inference still requires about 1.5TB of HBM bandwidth. To achieve profitable deployment while maintaining reasonable interaction speed, it is still necessary to connect a large number of chips through high-bandwidth networks such as GB300 NVL72 and rely on services like WideEP for scaling.
WideEP distributes 896 expert models to multiple GPUs, and performs Token distribution and result aggregation twice at each layer during each inference, requiring more than 120 inference steps per forward pass. In contrast, the KV cache transfer between prefilling and decoding occurs only once per dialogue round, so the bandwidth saved by KDA may be much less than the expanded network demand brought by large-scale expert models.
SemiAnalysis believes that a more efficient attention mechanism could also drive the increase of context length from 1 million Tokens to over 5 million Tokens. According to the Jevons Paradox, efficiency improvements may expand AI usage, thereby further increasing network requirements.
