header-langage
简体中文
繁體中文
English
Tiếng Việt
한국어
日本語
ภาษาไทย
Türkçe
Scan to Download the APP

Tencent conducted a blind test on the HY4, comparing it to the GLM 5.3 and the Kimi K3, with an output price 82% lower.

Mindseye Beating AI News Flash: Tencent also specifically released a set of real-world Hy4 preview engineering blind tests. 163 internal Tencent experts tested different models with 203 engineering tasks, with Hy4 ultimately averaging 2.99/4, higher than GLM-5.3's 2.92/4 and Kimi K3's 2.94/4.


Specifically, Hy4 had a win rate of 46.8% against GLM-5.3, a draw rate of 12.8%, and a loss rate of 40.4%; against Kimi K3, the win rate was 51.2%, draw rate 7.9%, and loss rate 40.9%. These questions are from real internal Tencent engineering tasks, not traditional fixed benchmarks.


However, Hy4 did not lead comprehensively in public benchmark tests. From the evaluation results released by Tencent, it had wins and losses against GLM-5.3, Kimi K3, and other models, with DeepSWE, CyberGym, and other code and network security tests still lagging behind GLM-5.3.


Pricing is the most obvious advantage of Hy4. It costs 6 yuan for every million Tokens input and 18 yuan for output, which is 25% cheaper than GLM-5.3 and about 36% cheaper; compared to Kimi K3, it is 70% and 82% cheaper, respectively. The cache hit is only 0.3 yuan, also 85% cheaper than GLM-5.3 and Kimi K3's 2 yuan.

举报 Correction/Report
Correction/Report
Submit
Add Library
Visible to myself only
Public
Save
Choose Library
Add Library
Cancel
Finish