Models & LLMs

Reka AI's Rho-1 Model Unifies Text, Images, Video, and Robot Control

Reka AI has unveiled a new 19-billion-parameter omni-model called Rho-1 that processes text, images, video, and robot control in a single neural network. The model is trained on 320 H100 GPUs in three months and uses a shared context window for all modalities.

The Decoder Β· Oct 05, 2026

What happened

  • Rho-1 is a 19-billion-parameter omni-model that processes multiple modalities in one neural network
  • Rho-1 is trained on 320 H100 GPUs over three months
  • Rho-1 processes all modalities as tokens in a shared context window

Why it matters

Reka AI's Rho-1 represents a significant advancement in multimodal AI, offering a unified approach to processing text, images, video, and robot control. This could lead to more efficient and integrated AI systems for various applications, from robotics to content generation.

The Elephant take

🐘 ιΌ‹ Rho-1 is a bold attempt to unify AI modalities, but the claim that it 'runs all modalities as tokens in one shared context window' sounds like a technical gimmick. The real test will be whether this model can handle real-world tasks without breaking down.

Who should care

  • AI researchers
  • Robotics developers
  • Multimodal AI enthusiasts

What to do next

  1. Evaluate the model's performance on real-world tasks
  2. Compare Rho-1 with existing multimodal models
  3. Check for independent verification of claims
  4. Stay updated on Reka AI's future releases

Keep in mind

The evidence is repetitive and lacks depth, suggesting a potential for misleading claims about the model's capabilities.

Read the original reporting at The Decoder β†—