Reka Launches Rho-1, a 19B Omni-Model Handling Text, Images, Video and Robot Actions in One Network
On Monday, October 5, Reka AI released a research preview of Rho-1 — a 19B-parameter omni-reasoning model, trained from scratch. Inside a single neural network it understands and generates text, images and video, and outputs robot actions directly: no agent pipeline, no tool calls, no second model. The pitch is blunt: collapse the dominant multi-model stack (a central model that plans, then delegates to specialists) into one — text, vision and robot actions unified as tokens in a single context window.

Brief · Oct 11
- 01
Amazon in Talks to Buy AI Startup Decart in Deal Valued Around $7 Billion
- 02
GSK Expands Chai Deal After Wet-Lab Validation of AI Designs
- 03
PPT Master Hits GitHub Trending: Documents Become Native PowerPoint
- 04
context-mode Hits GitHub Trending: Tool Output, Sandboxed First
What Happened
Reka AI — founded by ex-Google DeepMind senior scientist Dani Yogatama, which shipped the GPT-4-class Reka Core in April 2024 — launched Rho-1 via an official blog post on Monday, with THE DECODER and RuntimeWire covering the same day. The demo is a single five-turn session: draw a red-and-white lighthouse coast scene, box the lighthouse, animate it into a drone-approach video, restyle the weather into a snowstorm, then answer 'what changed between the two videos' — all inside the same model, the same KV-cache state, with zero tool calls and no second model.
Architecturally, Rho-1 is 'symmetric': inside every transformer block, two expert weight streams — one for understanding (language and visual parsing), one for generation (image/video denoising) — sharing attention and one KV cache; discrete tokens encode text, reasoning and commands, while continuous tokens encode image latents, video frames, robot actions and proprioception. Because inputs and outputs share the same formats, anything the model generates feeds straight back into context — predicting the next frame, the next instruction and the next action is, in Reka's words, 'one operation, not three models sharing a bus.'
Key Facts
- Speed numbers and the chart legend: The base model generates video at a median 0.79x real time, with a watchable stream starting in roughly 6 seconds; the distilled variant cuts the denoising trajectory from 99 steps to 8, rendering a 5.3-second clip in about a second — 'faster than any video model we timed,' the company says. The blog also shows a comparison: a 13.8-second multi-agent pipeline route versus Rho-1's measured 7.0 seconds. Per the legend, the pipeline bar is labeled 'illustrative' — only Rho-1's 7.0 seconds is measured.
- Training: 320 H100s, about 3 months: Rho-1 is a symmetric-architecture model trained from scratch. The blog acknowledges that multi-objective training (generation and understanding without dragging each other down) has no settled consensus and that training instability is an open research problem.
- Robot action data: the Inverse Dynamics Model: Robotics lacks action data — teleop logs are expensive and embodiment-specific. Reka addresses this with its Inverse Dynamics Model: ordinary internet video doesn't record steering angles or joint torques, so the IDM observes raw video, infers the underlying control signals, and feeds those actions to Rho-1 as native tokens. The demos show a banana lifted off a plate, a mug carried by its handle, and an apple picked from the wrist camera's view, under pure text conditioning.
- Reka's own Limitations section: The blog carries a standalone Limitations chapter: long-horizon drift (a 30-second stream keeps photorealistic texture while drifting into a structurally incompatible room layout), unreliable object grounding across video, brittle targeted editing across prompts, and native video capped at 672x384.
- Availability: research preview: This is a research preview — no released weights, no API, no third-party benchmarks. Every speed number is company-tested, and the demo scenes were chosen by the company.
Context
Google DeepMind has Genie 3 (text-generated, explorable real-time environments); Reka's Rho-1 puts simulation, reasoning and action into a single network. In September, Reward AI released OM-1, which trains robot policies on human demonstration data. Reka's previous flagship was the language model Reka Core (April 2024, benchmarked against GPT-4 / Claude 3 / Gemini Ultra); Rho-1 is its move into physical intelligence, using a symmetric architecture and an inverse-dynamics data pipeline.
Why it matters
Rho-1 understands and generates text, images and video and outputs robot actions inside one 19B-parameter network, with no tool calls; it is only a research preview, with no weights, API or third-party evaluation.
Comments
Today
Oct 12 Monday- BriefAmazon in Talks to Buy AI Startup Decart in Deal Valued Around $7 Billion
10 stories · Oct 11
- Quiz
- Call
Latest news
All →- Oct 11Amazon in Talks to Buy AI Startup Decart in Deal Valued Around $7 Billion
- Oct 11GSK Expands Chai Deal After Wet-Lab Validation of AI Designs
- Oct 11PPT Master Hits GitHub Trending: Documents Become Native PowerPoint
- Oct 11context-mode Hits GitHub Trending: Tool Output, Sandboxed First
- Oct 11Cloudflare Acquires Deno: Deploy Shuts Down in Six Months
- Oct 11Anthropic Updates Claude Usage Policy: Armed Drones Named and Banned, Effective Nov 12
