Loading video...

Video Failed to Load

Go Home

New model: your robot can now pack your suitcase 🧳 Xiaomi has released a new robot foundation model. Called Xiaomi-Robotics-1, it is designed to have a robot pick things up and move them around. But first, DEFINITIONS: - Mixture-of-Transformers (MoT): An architecture where separate transformer "experts" (e.g., one for...

15,777 views • 1 month ago •via X (Twitter)

22 Comments

PrismaX's profile picture
PrismaX1 month ago

👀

Léo's profile picture
Léo1 month ago

Working on this too 👀 ?

Daniel Lambert's profile picture
Daniel Lambert1 month ago

Foundation models that can generalize manipulation skills across everyday tasks are a big step toward more useful home robots. 🚀🤖

Léo's profile picture
Léo1 month ago

Yesss let's go 🦾

Praveen Kumar Verma's profile picture
Praveen Kumar Verma1 month ago

I have written about this few days back on it was open sourced. It is actually very good step for open sourcing the dataset where it can be used by other learners for training robots.

Léo's profile picture
Léo1 month ago

Definitely, and maybe a team with enough ressources can make use of the full dataset for training

Praveen Kumar Verma's profile picture
Praveen Kumar Verma1 month ago

Yes, with enough infrastructure a team can train robots precisely.

Léo's profile picture
Léo1 month ago

🦾

Tanay's profile picture
Tanay1 month ago

Which arms are those ?

Léo's profile picture
Léo1 month ago

I dont know actually, and it does not seem specified in the study either

Tanay's profile picture
Tanay1 month ago

China has a lot

Léo's profile picture
Léo1 month ago

Sure does

Zipper's profile picture
Zipper1 month ago

Very cool. I honestly think this kind of robot makes a lot more sense than humanoids.

Léo's profile picture
Léo1 month ago

it sure does, wheels make it easier to manage for the embedded model

Crimson_earth's profile picture
Crimson_earth1 month ago

Imagine paying $1000s only to pack your bag.

Léo's profile picture
Léo1 month ago

I would

Jeffrey Towson 陶迅's profile picture
Jeffrey Towson 陶迅1 month ago

Action chunking :)

Léo's profile picture
Léo1 month ago

Indeed!

Robotics Alpha's profile picture
Robotics Alpha1 month ago

the 20k vs 100k gap is the real story here. we saw same pattern at Tesla Optimus. Staged demos run on curated subsets, full corpus never surfaces. Xiaomi-Robotics-1's architecture is solid but that missing curve tells us where the bottleneck actually lives 🧳🤖

Léo's profile picture
Léo1 month ago

thanks gpt

TensorGamma's profile picture
TensorGamma1 month ago

10x speed haha

Léo's profile picture
Léo1 month ago

You have to do what you have to do 🫡

Related Videos

JUST IN: Dyna Robotics just published one of the most important research papers in robotics this year. It could fundamentally change how robot foundation models are trained. A scaling law that transfers from human video to robot performance. Dyna-2 is out and it's 🔥 Here's what that means in plain terms. Dyna-2 was pre-trained on ONE MILLION hours of egocentric human video, 170 years of continuous human experience, cooking, folding, assembling, cleaning. And as that human data scaled, robot performance improved. Predictably. Monotonically. Across 39 tasks on two different robot embodiments the model had never seen. → 1,000 hours pre-training → 20% normalised task performance → 10,000 hours → 28% → 100,000 hours → 45% → 1,000,000 hours → 53% Human video exists at effectively unlimited scale. Every cook, every factory worker, every craftsperson wearing a camera is generating training data for future robots. But the finding that stunned even the researchers, world modeling is what makes the transfer work. A model trained to predict future video AND actions massively outperforms one trained on actions alone. Video is the new scaling axis for robotics. One more jaw-dropping data point. 13 minutes of teleoperation data was enough to fine-tune Dyna-2 to open a bottle cap using two five-fingered robot hands. The robots are coming, and they're learning from us directly :D Read more here: Congrats Jason Ma and team! ~~ ♻️ Join the weekly robotics newsletter, and never miss any news →

Lukas Ziegler

23,681 views • 1 month ago

X Square Robot Unveils New Embodied AI Model, Says Robots Will Arrive in Homes in 35 Days Backed by Alibaba, ByteDance, Xiaomi and Meituan, X Square Robot unveiled a next-generation embodied AI foundation model for home robots and said its first deployments in everyday households will begin within 35 days. X Square Robot on Tuesday unveiled WALL-B, a new embodied AI foundation model designed for deployment in real-world homes, marking what the company described as a major step toward bringing general-purpose robots into daily family life. At a launch event themed "Born to Bot, Bot to Family," the company also introduced its World Unified Model (WUM) architecture, a training framework that combines vision, language, action and physical prediction within a single system from the outset. X Square said the model is intended to help robots operate in the far more unpredictable setting of a home, where tasks, layouts and interactions vary from moment to moment. "Robots in factories and in homes are completely different. In factories, they repeat the same action 10,000 times without variation. In a home, however, they need to perform 10,000 different actions, each unique and non-repetitive. Therefore, the challenge of a truly intelligent robot lies not in repeating a single action, but in the ability to execute new, untrained movements within unstructured environments. Deploying robots in the home is one of the most significant technical hurdles of our time," said Qian Wang, founder and CEO of X Square Robot. WALL-B is the first real-world implementation of the World Unified Model architecture. Unlike modular systems that train perception, language and control separately, X Square Robot said World Unified Model optimizes those capabilities jointly from the very beginning. The company said that allows physical prediction — including force, friction and collision dynamics — to emerge as part of the model itself, rather than being layered on afterward. "We train all capabilities—vision, language, action, and prediction—within the same network from day one. Much like infants, who do not learn to see, move and speak in isolated, sequential stages, but instead see, move listen and act simultaneously while receiving feedback, we have integrated all these capabilities into a unified whole," said Wang Hao, CTO of X Square. X Square Robot said the development of WALL-B rests on two pillars. The first is a data strategy that prioritizes training on authentic, non-staged home environments to cover the “long-tail” distribution of real-world scenarios, such as misplaced objects and temporary occlusions. Unlike models primarily trained on synthetic data or laboratory datasets, this strategy exposes WALL-B to the natural clutter of lived-in spaces—misplaced items, unexpected obstacles, and spontaneous human activity—ensuring that the training data reflects real-world conditions rather than a simplified version. The second is a physics-aware predictive mechanism that anticipates physical outcomes before an action is taken, enabling the model to respond to contact dynamics instead of just reacting. The development of the self-developed WUM architecture on physical robotic platforms highlights the company’s accumlated experience in bridging sim-to-real gaps across varied operational contexts. Wang commented that the current AI model is still in an "intern" stage, subject to errors requiring remote assistance. For instance, it may mistakenly place slippers in the kitchen or pause while wiping a table to "think". However, the model operates nonstop 24 hours a day, becoming increasingly "intelligent" as each day of operation generates new data. In 35 days, on May 25, X Square Robot will officially bring its robots into everyday homes, underscoring the company’s long-term commitment to the home robotics sector.

X Square Robot

52,968 views • 5 months ago

Excited to announce GR00T N1, the world’s first open foundation model for humanoid robots! We are on a mission to democratize Physical AI. The power of general robot brain, in the palm of your hand - with only 2B parameters, N1 learns from the most diverse physical action dataset ever compiled and punches above its weight: - Real humanoid teleoperation data. - Large-scale simulation data: we are open-sourcing 300K+ trajectories! - Neural trajectories: we apply SOTA video generation models to “hallucinate” new synthetic data that features accurate physics in pixels. Using Jensen’s words, “systematically infinite data”! - Latent actions: we develop novel algorithms to extract action tokens from in-the-wild human videos and neural generated videos. GR00T N1 is a single end-to-end neural net, from photons to actions: - Vision-Language Model (System 2) that interprets the physical world through vision and language instructions, enabling robots to reason about their environment and instructions, and plan the right actions. - Diffusion Transformer (System 1) that “renders” smooth and precise motor actions at 120 Hz, executing the latent plan made by System 2. We deploy N1 on GR1 robot, 1X Neo robot, and a large collection of simulation benchmarks. N1 achieves up to +30% boost in diverse manipulation tasks for household and industrial settings. While humanoid robots are the main focus of N1, our model also supports cross-embodiment. We finetune it to work on the $110 HuggingFace LeRobot SO100 robot arm! Open robot brain runs on open hardware. Sounds just right. Let’s solve robotics, together, one token at a time. Links to our Whitepaper, Github repo, HuggingFace model, and open dataset page in the thread: 🧵

Jim Fan

467,237 views • 1 year ago

The robot flipped a pancake nobody taught it! 🥞 Skild AI team assumed pancake flipping had to be somewhere in the training data. So they searched. Millions of hours of pre-training data. Nothing. S1 inferred the whole task from a single human demonstration. That's their new general robot model, built as an in-context learner from the ground up. Every new robot task today starts with days of teleoperation and a fine-tuning run on a specialist policy. S1 skips all of it. Much like a language model, it never updates its weights to learn a new task. The demonstration enters the context window, and the policy uses it to decide what to do next. → Ten-minute tasks it was never trained on, composed from primitives learned in pre-training: a new style of coffee, potting a plant, frying pancakes. → Soil and pots arrived at their office at 8:54 PM. The robot was running the task autonomously by 9:27 PM. → Slide objects away mid-reach, swap them, change the lighting, it still finishes. → The prompt waters a plant with a watering can, but only a cup is available. It uses the cup. It doesn't rigidly replay what it saw, but it recovers from its own errors, and sometimes executes with more precision than the demonstrator, when the human fumbles an egg and makes a mess, S1 performs the same step cleanly. The demonstration is a specification of the goal, and not a trajectory to copy. On unseen tasks after 100K hours of pre-training: language-prompted VLAs reach 9%. Their new model reaches 66%. It's already deploying with industrial partners, with a wider rollout over the coming months. Congrats Deepak Pathak and team behind this! 😮‍💨 🔗 Link to their latest blog: ~~ ♻️ Join the weekly robotics newsletter, and never miss any news →

Lukas Ziegler

11,406 views • 29 days ago

Chinese robotics company Astribot released their latest World-Action Model (WAM), Lumo-2. Technical breakdown: - based on a frozen 🥶 Qwen-3.5 4B VLM - trained in 3 progressive stages: 1. Action is aligned with latent world dynamics (an abstract representation of action). Real-world actions are anchored to physical constraints, while the latent space is guided to focus on motion-relevant changes. This bidirectional relationship makes the model physically grounded -> critical for a world model. 2. Action is aligned with vision and language. Reusing the vision backbone and action encoder from the frozen VLM, the authors add a custom vocabulary (for new actions), a semantic module, an action decoder, and an action projector. This aligns the (new) action representations with the (existing) vision-language semantic space. Most importantly: it builds a direct mapping from natural-language instructions to motor execution. 3. End-to-end training on language, video, and robot data. Only the new modules (everything outside the frozen backbone) are trained end-to-end across temporal reasoning, physical understanding, long-horizon, and dexterous manipulation. At the end of the day, Lumo-2 is not the best on benchmarks, but that's not the point. What's genuinely new: - a way to combine latent world modeling and action generation through progressive alignment - a physically-grounded latent dynamics space - it lifts performance on unseen objects using un-annotated human egocentric video + Vision Pro captures, no special transfer algorithm needed Why it matters: - the whole model is thin trainable adapters (semantic module, action decoder/projector) on a frozen 4B backbone (cheap) - that scale is suited for real-time embedded inference (~2.71× decode speedup, no accuracy loss) - its real moat is long-horizon execution, where the added temporal memory pays off far more than on any other task As a result, this robot can now make your latte (5x sped up video):

Léo

32,331 views • 2 months ago