BlockBeats News, September 9 — Ant Bailing officially open-sourced its first native multimodal model, Ling-3.0-flash-VL. The model had previously been available via API, and now BF16 and FP8 weights have been released under the MIT license, with FP4 and INT4 versions to follow.
Ling-3.0-flash-VL can read images, videos, documents, and software interfaces, then invoke tools to complete operations, and continue to inspect and refine based on execution results. For example, when building a webpage, it can write code based on a reference image, then review the actual rendered page and make adjustments on its own.
The model adopts a MoE architecture with approximately 124 billion total parameters, activating around 5.5 billion per inference. Its context window can be extended up to 1 million tokens.
Ling-3.0-flash-VL inherits the core strengths of Ling-3.0-flash as an efficient execution node in Agent workflows, balancing output quality and execution efficiency within a visual feedback loop, driving complete tasks forward at lower cost and in shorter time.

