The Maivia Gazette

Verified AI news, every morning

Research

Reka AI previews Rho-1, a 19-billion-parameter model that handles text, images, video and robot control in one network

The same weights that predict camera images also drive robot movements, and its robot control was learned from ordinary internet videos.

A robotic arm waters seedlings in a greenhouse beside floating film frames of the same motion.
AI-generated illustration, not event photography. The motion is AI-generated from the still.

Reka AI has released a research preview of Rho-1, a 19-billion-parameter omni-model that processes and generates text, images, video and robot control actions in a single neural network. Most AI systems route tasks to specialized models. Rho-1 instead treats every modality as tokens in one shared context window, with no tool calls or external models. Reka says the model generates continuous video in real time and can respond to new instructions on the fly without restarting. The same weights that predict camera images also drive robot movements. Robot training data is scarce, so Reka built an inverse dynamics model that extracts control signals from ordinary internet videos and used them to train Rho-1. The model was trained on 320 H100 GPUs over about three months. Reka previously released Reka Core, a multimodal language model, in April 2024. Rho-1 is part of a broader research push toward world models, which aim to connect perception, prediction and physical action. Its approach could ease the shortage of robot data if signals learned from web video carry over well to real robots, something the preview does not yet establish.

Sources

  1. The DecoderReka AI's omni-model Rho-1 handles text, images, video, and robot control in a single modelPublished · fetched

Also in this edition