Reka Releases Rho-1 to Combine Video Creation and Robot Actions in One Model
The research preview keeps generated scenes and later instructions in the same model state. Reka’s speed claims come from internal tests, and its playback demo shortens generation waits.
Listen to this story
The audio brief
Story brief
3 key pointsRho-1’s distinguishing claim is that video edits, scene interpretation and action outputs can draw on one evolving internal state: it marks objects and explains edits without a separate detector or captioning model. Released October 5, 2026, the 19-billion-parameter model accepts instructions while video is being generated and uses the same weights for image prediction and robot movement. Reka reports that Rho-1 Flash cuts denoising from 99 steps to eight; shown 5.3-second clips took 1.11–2.01 seconds. The training run reportedly used 320 H100 GPUs for about three months, and the performance figures are company-reported.
- 01
Rho-1 uses understanding and visual-generation streams with shared attention and a KV cache to carry context through a session, according to Reka.
- 02
In an aerial-scene demo, “bank left” and “bank right” continuations share the opening half-second before diverging.
- 03
Reka says an inverse-dynamics model derives robot controls from ordinary internet videos to help address scarce robot-training data; The Decoder reported the training scale.
Creating a scene, turning it into video and asking what changed can now happen inside one model in Reka’s research preview. Released October 5, 2026, Rho-1 is a 19-billion-parameter system that the company says combines text, images, video and robot actions in a shared context, rather than passing work between separate models.
The lighthouse stays inside the conversation
Reka’s five-turn demonstration starts with a red-and-white lighthouse on a rocky coast. The user asks the model to mark the lighthouse, animate a drone’s approach, replace the weather with a snowstorm and explain the difference between the videos. Each response uses the same model and accumulated conversation state.
Underneath that workflow are two connected streams: one handles understanding and text responses, while the other generates visual material. Both use shared attention and a common KV cache—the stored internal information carried forward through the session. Reka says there is no tool call or second model producing those responses.
The distinction is more than keeping a chat history. When Rho-1 animates the lighthouse, it uses the image representation already in context rather than re-encoding an exported picture. When it explains an edit, it reads the internal representation that produced the change, rather than sending finished frames through a separate captioning model.
Object marking takes a different path through that architecture. The model outputs coordinates for the box using the scene already in its context, without a separate object detector. Reka also says the model itself plans the shot before rendering it, rather than relying on an added prompt-enhancement system.
Flash reduces the work between prompt and pixels
Reka reports that base Rho-1 generates video at a median 0.79 times real time, with a watchable stream beginning after roughly six seconds. A faster, distilled variant called Rho-1 Flash reduces denoising—the repeated steps that turn a noisy representation into an image or clip—with what Reka describes as minimal quality loss.
In the company’s shown Flash examples, 5.3-second clips render in 1.11–2.01 seconds as prompts change a dog’s breed, clothing and surroundings. These are company-reported timings. Reka’s broader speed comparisons come from internal testing; its pipeline comparison is illustrative, and the lighthouse playback shortens generation waits.
Instructions can arrive while the world is moving
Rho-1 also supports continuous generated video that can take new instructions without restarting, according to Reka. One demonstration branches an aerial scene into “bank left” and “bank right” continuations. Both share the opening half-second before following different paths. Another shows a simulated robot arm picking up and putting down a ball as commands arrive.
During a live rollout, a new command enters the understanding stream and updates the stored state. The generation stream then renders the changed trajectory. Reka proposes using such reactive environments to test autonomous-driving systems or train robot-control policies, with developers changing conditions while a simulation runs.
Robot control extends the shared-model idea beyond generating footage. Reka says the same model weights that predict camera images also drive robot movements. To address scarce robot training data, the company built an inverse dynamics model that derives control signals from ordinary internet videos, The Decoder reports.
Reka trained Rho-1 from scratch. The training run used 320 H100 GPUs over about three months, according to The Decoder, covering the model that combines media generation and understanding with action output.
Sources
- the-decoder.comReka AI's omni-model Rho-1 handles text, images, video, and robot control in a single model
- reka.aiRho-1: Collapsing the multimodal stack
Reader comments
Newest comments first. Replies stay oldest first.