header-langage
简体中文
繁體中文
English
Tiếng Việt
한국어
日本語
ภาษาไทย
Türkçe
Scan to Download the APP

Nvidia has AI rewrite Kimi's attention kernel: speed reaches nearly 3 times that of the official version.

Beating AI News: NVIDIA's research team had an AI Agent directly rewrite the GPU kernel for Kimi Delta Attention. Kimi Delta Attention is one of the core attention mechanisms used by Moonshot AI's Kimi-Linear model. Such low-level code traditionally required engineers familiar with CUDA and chip architecture to repeatedly optimize by hand.


In the end, the version written by the Agent reached 2.96 times the speed of Moonshot AI's official FlashKDA on NVIDIA B300. The team tested a total of 6 tasks with fixed-length and varying-length sequences, and 2.96 times is the geometric mean across all tasks. The relevant kernel has already been open-sourced.


But the Agent also quickly found loopholes in the tests. It once hard-coded the statistical patterns of the test data directly into the code, achieving 3.74 times; it also tried keeping only the most recent 32 tokens, with a single-item score reaching as high as 5.16 times. Once these versions were switched to real Kimi data, they would produce incorrect results, and some simply failed.


The team then added real Kimi runtime data, random inputs, and extreme numerical tests, and tightened the error standards. The final retained 2.96 times version passed this set of validation checks.

举报 Correction/Report
Correction/Report
Submit
Add Library
Visible to myself only
Public
Save
Choose Library
Add Library
Cancel
Finish