header-langage
简体中文
繁體中文
English
Tiếng Việt
한국어
日本語
ภาษาไทย
Türkçe
Scan to Download the APP

DeepSeek Visual Model Multimodal Agent approaches Opus 4.8! 4 hits in 2:2 format

Dynamic Beating AI News Flash: DeepSeek has released the first batch of Agent scores for V4-Flash-Vision-Exp. In the 4-mode Agent evaluation, it tied 2-2 with Opus 4.8: Agents' Last Exam was 27.3 to 25.7, ZeroBench was 35.0 to 34.0; ApexBench was 36.5 to 39.4, and Chartography was 64.3 to 65.0.

Compared to the plain text version V4-Flash-0731, the Vision version increased from 26.2 to 36.5 in ApexBench, and from 25.2 to 27.3 in Agents' Last Exam. However, these two evaluations themselves include images. The official statement specifies that V4-Flash will directly ignore the multimodal content. Therefore, this primarily reflects the improvement in ability with the addition of visual input rather than a sudden enhancement in pure text reasoning.

With the addition of vision, the text Agent's ability did not show a significant reduction. In the 7 text evaluations, Vision Exp scored higher in 6 items compared to V4-Flash-0731. For example, DeepSWE increased from 54.4 to 59.3, even surpassing Opus 4.8 at 58.0; Toolathlon scored 75.9, almost equaling Opus 4.8 at 76.2.

It is important to note that these results are from DeepSeek's official self-evaluation and not from a third-party independent ranking. The DeepSeek series uses the DeepSeek Harness in ultra-light mode for public Code Agent text tasks, with the reasoning level set to max.

举报 Correction/Report
Correction/Report
Submit
Add Library
Visible to myself only
Public
Save
Choose Library
Add Library
Cancel
Finish