According to Perceive Beating monitoring, Artificial Analysis has finally released the standalone benchmark score for Qwen3.8-27B: an Intelligence Index of 52 points.
This Dense model with only 27B parameters directly matches the performance of DeepSeek V4 Flash 0731 and GPT-5.6 Luna Max, falling just 1 point short of DeepSeek V4 Pro 0813 and GLM-5.2, which scored 53 points. The previous generation, Qwen3.6-27B, scored only 38 points.
Even more remarkable is the Agent capability. Qwen3.8-27B's Agentic Index has reached 51 points, surpassing DeepSeek V4 Pro with 50 points, DeepSeek V4 Flash with 48 points, and GPT-5.6 Luna with 47 points. This metric specifically assesses the model's ability to invoke tools, plan steps, and consistently complete complex tasks.
However, achieving 52 points comes at a cost. When Artificial Analysis ran the full Intelligence Index suite, Qwen3.8-27B generated approximately 160 million output Tokens, clearly falling into the category of "intelligence stacked with tokens."
The Qwen official default reasoning level is set to xhigh, while also supporting medium, low, and no thinking modes. Currently, Artificial Analysis has not specified which reasoning level was used, so it cannot be confirmed whether this performance is based on the default xhigh setting.

