Beating AI News Flash: The open-source Computer Use tool Cua has added an optional Perception visual extension to Cua Driver. Driver originally prioritized finding buttons through webpage structure and system accessibility trees. When encountering Canvas, games, remote desktops, or custom-drawn applications, these structures may be empty. Perception directly reads screenshots, breaking the image into regions with text, types, and positions for the Agent to continue operating.
The first version uses OmniParser to recognize icons and controls, and PP-OCR to read text. The entire parsing process is completed on the local CPU, requires no GPU, and does not send screenshots to the network. Official tests show that a single screenshot takes about 2 to 2.5 seconds on macOS, 3.5 to 4 seconds on Linux, and 8 to 9 seconds on Windows. It mainly serves as a fallback when page structure and accessibility trees fail, and does not run visual parsing by default at every step.
Cua also restricts the reuse of old screenshots for visual operations. Each visual recognition is bound to the current screenshot's `capture_id`, and subsequent clicks must come from the same screenshot. If the screenshot exceeds 60 seconds, or if an operation has already been performed once, Driver will refuse to continue using it, avoiding mistaken clicks with old coordinates after the interface changes.
Cua could actually already perform visual localization through `cua-som` and OmniParser. Now this capability has been directly integrated into Cua Driver, and upper-layer Agents can obtain recognition results through a unified interface. Even if the model itself cannot see images, it can decide the next operation based on this text and region information.
Perception is not a default component, and its OmniParser detector uses AGPL-3.0; Cua Driver remains under the MIT license.

