header-langage
简体中文
繁體中文
English
Tiếng Việt
한국어
日本語
ภาษาไทย
Türkçe
Scan to Download the APP

Thousand Realms' $2.7B Hit to Opus 4.6 Door: Muse Glimmer Records 8 Consecutive Losses

According to Watchful AI monitoring, Qianwen has officially open-sourced Qwen3.8-27B, and the official benchmark scores have also been released. This local model with only 27B parameters has surpassed Meta's recently released 30B model, Muse Glimmer, in all 8 directly comparable tests listed by the official source.

The difference is particularly noticeable in Agent and Coding. In Terminal-Bench 2.1, Qwen3.8-27B scored 73.0, compared to Muse Glimmer's 51.7; in SWE-bench Pro, it was 61.7 versus 51.2; and in OSWorld-Verified, the scores were 84.3 versus 65.9. Qwen3.8-27B is also ahead in general inference, documentation, and visual tests.

Even more remarkable is the comparison with Claude Opus 4.6. Out of the 19 tests where both models achieved scores, Qwen3.8-27B won in 15. It has already surpassed SWE-bench Pro, LiveCodeBench, OSWorld, AndroidWorld, and multiple visual tasks; however, it still lags behind in Terminal-Bench, GPQA, HLE, and NL2Repo.

For a 27B model that can run locally on a personal computer, the official benchmark scores have reached the level of the previous generation's closed-source flagship.

举报 Correction/Report
Correction/Report
Submit
Add Library
Visible to myself only
Public
Save
Choose Library
Add Library
Cancel
Finish