Beating AI News Flash: The Databricks team says they put GPT-6 Astra and Opus 5 into an automated optimization loop, having the models continuously rewrite, test, and improve GPU kernels, ultimately ranking first across all four tracks of NVIDIA SOL-ExecBench. This benchmark includes 235 GPU kernel tasks, covering basic operators, complex fused operators, low-precision computation, and real large-model inference. The team says the model token cost for the entire process was about $70,000.
This system is built on KDA and Humanize, allowing AI to write kernels itself, run correctness and performance tests, and then continuously revise them based on the results. The KDA team recently also applied a similar method to Kimi Delta Attention, achieving 2.96 times the speed of Moonshot AI's official FlashKDA on NVIDIA B300.
The two achievements belong to different applications of the same Agent kernel optimization system. Kimi demonstrated that a highly difficult attention kernel can be rewritten by AI; this time, Databricks extended this method to hundreds of different kernels.

