header-langage
简体中文
繁體中文
English
Tiếng Việt
한국어
日本語
ภาษาไทย
Türkçe
Scan to Download the APP

DeepSeek is finally getting ready for multimodality? The brand-new vision model has been integrated into the official Harness

Perceive Beating AI News Flash: A new vision model, deepseek-v4-flash-vision-exp, has surfaced in DeepSeek. Some community members have successfully run it through the API, enabling direct image input for inquiries. A clear sign came today as DeepSeek Harness officially added this model to the default model directory, explicitly marked to support image input. The corresponding code PR is titled "publish the vision model." DeepSeek has not made an official announcement yet, and the model has not appeared in the official API documentation.

DeepSeek has actually been capable of image recognition for some time now. Image recognition mode has been deployed on both the website and app. Previous research from DeepSeek's vision team was based on V4-Flash, allowing the model to insert points and bounding boxes directly into the inference chain, essentially "pointing" at the image while conducting visual CoT inference.

This release had been foreshadowed. Just two days ago, DeepSeek Harness RC.8 had just added native image input to the official DeepSeek adapter. The release notes at the time already referenced deepseek-v4-flash-vision-exp, purposely not included in the default model directory to wait for the model endpoints to be ready.

Two days later, this restriction was formally lifted. Harness v0.1.1-rc.1 has now listed deepseek-v4-flash-vision-exp as the default vision model. Initially on hold for "endpoint readiness," it has now directly entered the official model roster, signaling an imminent release of the new vision model.

举报 Correction/Report
Correction/Report
Submit
Add Library
Visible to myself only
Public
Save
Choose Library
Add Library
Cancel
Finish