header-langage
简体中文
繁體中文
English
Tiếng Việt
한국어
日本語
ภาษาไทย
Türkçe
Scan to Download the APP

Reka releases 19B Rho-1: a single model that simultaneously handles images, text, video, and robot actions.

动察 Beating AI News: Reka, a US AI company founded by former researchers from DeepMind, Google Brain, and Meta FAIR, has released a 19-billion-parameter model called Rho-1. It can handle text, images, and video in a single conversation, and can also output robot actions. For example, it can first generate an image of a lighthouse, then make the lighthouse move, modify the weather in the video, and finally continue by asking what the difference is between two videos. Previously generated content remains in the context the whole time and does not need to be handed off to another model for processing.


It also supports continuing to receive instructions during video generation. While a video is being generated forward, the user can temporarily ask to change the weather, the direction of movement, or other content, and the later frames will adjust accordingly. The base version generates at about 0.79x real time. The distilled version compresses the generation steps from 99 to 8, and Reka measured that it can generate about 5.3 seconds of video in about 1 second.


Robotics is also connected within the same model. Rho-1 predicts what the camera will see next while outputting robot action signals, without needing a separate planning model to decide how to move. At present, the official demonstration is still limited to LIBERO simulation tasks, and no test results on physical robots have been released.


Similar unified models have already appeared this year. NVIDIA Cosmos 3 can also uniformly process text, images, video, audio, and actions. Rho-1 leans more toward continuous interaction, with emphasis on repeatedly generating, modifying, and continuing reasoning within the same context.


Rho-1 was trained from scratch using 320 H100s for about 3 months. The current video resolution is 672×384, and long videos may still show structural drift in the image, while video grounding and local editing are also not stable enough. The model is currently only a research preview, and neither the weights nor the API have been released.

举报 Correction/Report
Correction/Report
Submit
Add Library
Visible to myself only
Public
Save
Choose Library
Add Library
Cancel
Finish