Xiaomi Open-Sources Xiaomi-Robotics-U0: First Unified Generative Model to Master Four Embodied AI Tasks
Xiaomi releases a 38-billion-parameter multimodal foundation model that unifies scene generation, trajectory migration, video synthesis, and image editing for robotics, ranking first on the WorldArena benchmark while delivering 83x inference acceleration.
Xiaomi has officially open-sourced Xiaomi-Robotics-U0, a 38-billion-parameter multimodal autoregressive foundation model designed to serve as a universal generative engine for embodied intelligence. Released on July 15, the model represents the industry's first unified architecture capable of handling four distinct robotic tasks—scene generation, trajectory transfer, interaction video synthesis, and general image editing—within a single framework.
Four Capabilities, One Model
Unlike conventional robotics models that specialize in narrow domains, Xiaomi-Robotics-U0 integrates four critical pipelines for synthetic data generation and manipulation:
Embodied Scene Generation enables the model to synthesize multi-view initial environments from text descriptions tailored to specific robot embodiments. Whether simulating desktop workstations, kitchen environments, warehouse floors, or complex open-world settings, the system generates physically plausible observations that match linguistic prompts.
Embodied Transfer allows existing robot trajectories to be seamlessly migrated into novel contexts. The model can alter lighting conditions, backgrounds, surface materials, target objects, and workspace aesthetics while preserving the original mechanical arm poses and spatial layouts—effectively multiplying training data diversity without physical recollection.
Robot Interaction Video Generation produces temporally coherent video sequences based on initial visual observations and action commands. The system maintains physical consistency and motion continuity across frames, demonstrating zero-shot generalization to previously unseen scenarios.
General Visual Synthesis retains robust text-to-image and anything-to-image capabilities, enabling knowledge transfer from vast internet visual data into embodied intelligence applications.
Performance and Acceleration
Through the proprietary FlashAR+ inference acceleration architecture, Xiaomi-Robotics-U0 achieves an 83-fold speed improvement over standard autoregressive generation paradigms, addressing a critical bottleneck for industrial deployment.
On the WorldArena evaluation benchmark—a comprehensive test involving 126 global models—the system secured first place overall (competing under the anonymous identifier "UNIS"). Real-world validation demonstrated substantial practical benefits: when training policies on data augmented by Xiaomi-Robotics-U0, task completion rates improved by over 26% in out-of-distribution conditions featuring unfamiliar lighting and novel backgrounds.
Addressing Data Scarcity
The model offers a scalable solution to the chronic data shortage in robotics training. It can generate dangerous, extreme, or long-tail scenarios that are difficult or impossible to capture with physical hardware, while simultaneously enriching existing datasets through geometrically consistent augmentations—swapping objects, varying illumination, altering backgrounds, or introducing visual distractions without requiring new physical demonstrations.
Open Source Availability
Xiaomi has released the complete codebase and model weights without restriction. Developers can access the project through the following channels:
- Project Homepage: robotics.xiaomi.com/xiaomi-robotics-u0.html
- Code Repository: github.com/XiaomiRobotics/Xiaomi-Robotics-U0
- Model Weights: huggingface.co/collections/XiaomiRobotics/xiaomi-robotics-u0
- ModelScope: modelscope.cn/collections/XiaomiRobotics