Model Releases
Xiaomi-Robotics-1: New robotics model released
Xiaomi-Robotics-1 is a robot foundation model trained on over 100K hours of real-world manipulation trajectories. It is a Vision-Language-Action (VLA) model engineered for out-of-the-box mobile manipu
Xiaomi-Robotics-1 is a robot foundation model trained on over 100K hours of real-world manipulation trajectories. It is a Vision-Language-Action (VLA) model engineered for out-of-the-box mobile manipulation in unseen environments and efficient adaptation to new tasks. XR-1 follows a two-stage training paradigm inspired by large language models — pre-training for breadth, followed by post-training for alignment. It showcases that pre-training scaling behavior reliably transfers through post-training to real-world robot performance, with no signs of saturation. XR-1 couples a pre-trained VLM (Qwen3-VL) with a Diffusion-Transformer (DiT) via a Mixture-of-Transformers (MoT) — the DiT matches the VLM in layer count but uses a smaller hidden size for faster inference. HugginFace: https://huggingface.co/collections/XiaomiRobotics/xiaomi-robotics-1 GitHub: https://github.com/XiaomiRobotics/Xiaomi-Robotics-1 Paper: https://arxiv.org/abs/2607.15330 submitted by /u/121507090301 [link] [comments]
Related
- Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories
- See What Matters: Differentiable Grid Sample Pruning for Generalizable Vision-Language-Action Model
- SafeVLA: Towards Safety Alignment of Vision-Language-Action Model via Constrained Learning
Source: r/LocalLLaMA | 2026-08-05