Reka AI previews Rho-1, a 19-billion-parameter model that handles text, images, video and robot control in one network
The same weights that predict camera images also drive robot movements, and its robot control was learned from ordinary internet videos.

Reka AI has released a research preview of Rho-1, a 19-billion-parameter omni-model that processes and generates text, images, video and robot control actions in a single neural network. Most AI systems route tasks to specialized models. Rho-1 instead treats every modality as tokens in one shared context window, with no tool calls or external models. Reka says the model generates continuous video in real time and can respond to new instructions on the fly without restarting. The same weights that predict camera images also drive robot movements. Robot training data is scarce, so Reka built an inverse dynamics model that extracts control signals from ordinary internet videos and used them to train Rho-1. The model was trained on 320 H100 GPUs over about three months. Reka previously released Reka Core, a multimodal language model, in April 2024. Rho-1 is part of a broader research push toward world models, which aim to connect perception, prediction and physical action. Its approach could ease the shortage of robot data if signals learned from web video carry over well to real robots, something the preview does not yet establish.