header-langage
简体中文
繁體中文
English
Tiếng Việt
한국어
日本語
ภาษาไทย
Türkçe
Scan to Download the APP

SemiAnalysis: Kimi K3 Significantly Reduces KV Transmission Bandwidth, but Does Not Reduce AI Network Demand

BlockBeats News, July 19th - Semiconductor and AI research firm SemiAnalysis published an article stating that although about three-quarters of Kimi K3's network layer adopts KDA, which can reduce KV cache transfer bandwidth by up to 10 times compared to a full-scale global attention model, this does not mean that the AI network switch market will shrink significantly.


Kimi K3 has 2.8 trillion parameters, and even with MXFP4, each inference still requires about 1.5TB of HBM bandwidth. To achieve profitable deployment while maintaining reasonable interaction speed, it is still necessary to connect a large number of chips through high-bandwidth networks such as GB300 NVL72 and rely on services like WideEP for scaling.


WideEP distributes 896 expert models to multiple GPUs, and performs Token distribution and result aggregation twice at each layer during each inference, requiring more than 120 inference steps per forward pass. In contrast, the KV cache transfer between prefilling and decoding occurs only once per dialogue round, so the bandwidth saved by KDA may be much less than the expanded network demand brought by large-scale expert models.


SemiAnalysis believes that a more efficient attention mechanism could also drive the increase of context length from 1 million Tokens to over 5 million Tokens. According to the Jevons Paradox, efficiency improvements may expand AI usage, thereby further increasing network requirements.

举报 Correction/Report
Correction/Report
Submit
Add Library
Visible to myself only
Public
Save
Choose Library
Add Library
Cancel
Finish