动察 Beating AI News: Reka, a US AI company founded by former researchers from DeepMind, Google Brain, and Meta FAIR, has released a 19-billion-parameter model called Rho-1. It can handle text, images, and video in a single conversation, and can also output robot actions. For example, it can first generate an image of a lighthouse, then make the lighthouse move, modify the weather in the video, and finally continue by asking what the difference is between two videos. Previously generated content remains in the context the whole time and does not need to be handed off to another model for processing.
It also supports continuing to receive instructions during video generation. While a video is being generated forward, the user can temporarily ask to change the weather, the direction of movement, or other content, and the later frames will adjust accordingly. The base version generates at about 0.79x real time. The distilled version compresses the generation steps from 99 to 8, and Reka measured that it can generate about 5.3 seconds of video in about 1 second.
Robotics is also connected within the same model. Rho-1 predicts what the camera will see next while outputting robot action signals, without needing a separate planning model to decide how to move. At present, the official demonstration is still limited to LIBERO simulation tasks, and no test results on physical robots have been released.
Similar unified models have already appeared this year. NVIDIA Cosmos 3 can also uniformly process text, images, video, audio, and actions. Rho-1 leans more toward continuous interaction, with emphasis on repeatedly generating, modifying, and continuing reasoning within the same context.
Rho-1 was trained from scratch using 320 H100s for about 3 months. The current video resolution is 672×384, and long videos may still show structural drift in the image, while video grounding and local editing are also not stable enough. The model is currently only a research preview, and neither the weights nor the API have been released.

