Reka Rho-1: 19B omni-reasoning model for text, video and robot actions
Research preview of a 19-billion-parameter omni model that understands and generates text, images and video, and emits robot actions, all in one network.
Why it matters
Collapses the agentic pipeline into a single model: instead of a central model delegating to modality specialists, Rho-1 keeps text, vision and robotic actions as tokens in one shared context window, with continuous real-time video generation that can be steered mid-rollout. It trained on 320 H100s for about 3 months.
What you could build with it
An indie developer or robotics startup could use a single omni model like this to prototype a warehouse demo in which the same model narrates what it sees, renders a preview video of the planned motion, and then drives the arm - one model replacing a chain of specialists.
Does it hold up?
Too early to judge - research preview only; no public weights, API or pricing yet, so there is no third-party hands-on evidence of how it performs outside Reka's demos.
Built with Reka Rho-1
- Reka Releases Rho-1: A 19B Omni-Reasoning Model (MarkTechPost analysis with comparison table)article · In-depth breakdown of Rho-1's architecture (two expert streams, shared KV cache, 99-to-8-step distillation) and how it compares to ByteDance BAGEL, BAAI Emu3.5 and Google Genie 3.
- Reka AI's omni-model Rho-1 handles text, images, video, and robot control in a single model (The Decoder)article · Coverage of the release, noting the inverse-dynamics model that extracts control signals from ordinary internet video to work around scarce robot training data.
- Reka releases a 19B model for video generation and robot control (Runtime Wire, with announcement video)article · Walkthrough of Reka's X announcement demo: a lighthouse image located, animated, weather-changed and questioned about - all in one shared state with no tool calls.
Learn more
First spotted on article: source.
More AI models
ElevenLabs Eleven v4 + v4 Turbo: new emotive TTS architecture with a 100ms real-time variantElevenLabs released Eleven v4 and Eleven v4 Turbo on September 28: a new text-to-speech architecture with inline…models · JEV 0.72VoiceStudio: open-source, fully-local ElevenLabs alternative with voice cloning, dubbing and dictation in 646 languagesVoiceStudio is an open-source, fully-local voice AI studio: voice cloning from a 3-second sample, voice design, video…models · JEV 0.71Mistral Forge: enterprise platform for training and continuously improving proprietary modelsMistral AI launched Forge (Sep 28) — a full-lifecycle model training platform (pre-training, SFT, DPO/ODPO, RL…models · JEV 0.69DeepSeek V4-Flash official API: public beta with upgraded agent capabilities, Responses API and Codex supportDeepSeek's official V4-Flash is now a public beta API with massively upgraded agent capabilities - benchmark scores…models · JEV 0.69kev: open, trainable Jev-like family of small decision models on Qwen3.5/3.8A family of small decision models (0.8B to 27B) built on Qwen3.5/3.8 that reproduces Jev's prefill-only typed-decision…models · JEV 0.68GPT-Synopsys: OpenAI and Synopsys multi-year deal to build an AI model that operates EDA toolsOpenAI and Synopsys signed a multi-year agreement on Sep 30 to jointly develop GPT-Synopsys, a specialized model that…models · JEV 0.66
Get the week's best AI launches, plus 3 ideas worth building
One email every Saturday. Ranked by traction, not hype. Free.