How can we more effectively leverage robot data from... different embodiments for skill transfer? Excited to share that our new work, RoVi-Aug, has been accepted to Conference on Robot Learning as an oral paper! WIth RoVi-Aug, you can augment an existing robot dataset into a different robot and different viewpoints. A policy trained on the augmented dataset can zero-shot deploy on the unseen target robot with significantly different camera angles! 🧵👇 🔗 Check out our paper:show more

Chenfeng_X
28,321 次观看 • 2 年前
How can we leverage diverse human videos to improve... robot manipulation? Excited to introduce EgoVLA — a Vision-Language-Action model trained on egocentric human videos by explicitly modeling wrist & hand motion. We build a shared action space between humans and robots, enabling seamless transfer. With some robot demos, EgoVLA becomes a powerful, generalizable robot policy.show more

Ruihan Yang
58,800 次观看 • 1 年前
🚀Our New Paper on Open-Source Bipeda Robot MEVITA is... out! All components can be procured through e-commerce, and the robot is built with a minimal number of parts. All hardware, software, and learning environment are released as open source. 🌐 Thread👇show more

Kento Kawaharazuka / 河原塚 健人
48,768 次观看 • 1 年前
Can we make Transformers better and more efficient for... robot learning? Excited to introduce Body Transformer (BoT), an architecture that leverages robot embodiment in the attention mechanism, by treating it as a graph of sensors and actuators.show more

Carlo Sferrazza
60,679 次观看 • 2 年前
What's different between these two BC policies? It's the... same architecture, training budget, and data collection setup — the only difference is the controller gains! Controller gains are an understudied design parameter in robot learning. In our new work (w/ Antonia Bronars*, Pulkit Agrawal), we show how they act as an inductive bias across BC, RL, and Sim2Real transfer, with real consequences on performance. Here's what we found 🧵 * Equal Contribution 📄arxiv: 🔗website:show more

Younghyo Park
161,481 次观看 • 5 个月前
𝗥𝗼𝗯𝗼𝘁𝘀 𝗱𝗼𝗻’𝘁 𝗻𝗲𝗲𝗱 𝗺𝗼𝗿𝗲 𝗱𝗲𝗺𝗼𝗻𝘀𝘁𝗿𝗮𝘁𝗶𝗼𝗻𝘀. 𝗧𝗵𝗲𝘆 𝗻𝗲𝗲𝗱 𝘁𝗼 𝗹𝗲𝗮𝗿𝗻... 𝗳𝗿𝗼𝗺 𝗳𝗮𝗶𝗹𝘂𝗿𝗲 — 𝗮𝗳𝘁𝗲𝗿 𝘄𝗮𝘁𝗰𝗵𝗶𝗻𝗴 𝗵𝘂𝗺𝗮𝗻𝘀. Most robot learning systems assume failure is the end of learning. In our new work, we study whether robots can improve after deployment by learning from their own failures, without any human intervention, teleoperation, or corrective labels. The key idea is simple: human videos contain structure about how the world works. We use them to learn cross-embodiment representations of action, dynamics, and value, enabling a shared predictive space between human behavior and robot experience. This allows a new learning loop: 👉 pretrain on human videos 👉 deploy robot policy 👉 observe failures 👉 reinterpret failures using human priors 👉 improve autonomously We evaluate this across 7 real-world manipulation tasks, showing: 📈 40% → 81% success rate 🏆 Strong improvements over π0.6 RECAP and RISE ✔️ Zero human intervention during post-deployment improvement 🧬 Generalizes across robot embodiments and policy backbones A key finding is that explicit failure repair significantly outperforms failure reweighting, yielding substantially larger gains under identical data conditions (+25 pts vs +5 pts on the same π0.5 base policy). Overall, the results suggest a shift in how we think about robot learning: Human videos are not only for pretraining policies. They can provide the structure needed for continual self-improvement after deployment. 📄 Paper: 🌐 Project: I am grateful for working with the fantastic leads Hanzhi Chen and Anran Zhang, and our collaborators Simon Schaefer, Kejia Chen, Shi Chen, Daniel Cremers. Special thanks to Stefan Leutenegger for co-advising this project with me. ETH Zürich TU München Microsoft Check out Hanzhi's 🧵 for more detailsshow more

Oier Mees
12,379 次观看 • 2 个月前
Hey #NeuraxonMini is literally out! , we manage to... "transplant" a Neuraxon 2 bioinspired #AI brain to a physical robot the #SpheroMini moving from our last Scientific Paper (link bellow) by David Vivancos - e/acc & Jose Sánchez for Qubic #OpenScience hybridized with #Aigarth to the real World. First you need a Sphero Education Mini robot about 50$ Then you can try the first cool demos at Hugging Face: 1.- Neuraxon2MiniControl to drive the sphero robot 2.- Neuraxon2MiniWrite to write letters or words with physical moves of the sphero robot using Neuraxon Video Tutorials on youtube later today. Why this matters? Remember we are not building "dead" LLMs we are building #AliveAIs and for that we need to explore how it behaves in reality, from how it learns to how it fails, and what better way that in the emerging field of #robotics , time will tell if your next #HumanoidRobot have a #Neuraxon brain... Read the Paper: Explore the Neuraxon code here: Are you ready for #TrueAI ?show more

David Vivancos - e/acc
29,293 次观看 • 6 个月前
we open sourced the code to transform human videos... into robot trajectories, so you can train robots with your hands 👐🏻 we used it in our recent paper R+X: Retrieval and Execution from Everyday Human Videos (ICRA 2025 🇺🇸) link and details in thread 🧵show more

Norman Di Palo
10,604 次观看 • 1 年前
Excited to share a few presentations, demos, and workshop... talks from our group and collaborators at #ICRA2026! We will present recent work on real-to-sim-to-real robot policy evaluation, model-based planning with learned dynamics, and multi-modal manipulation. We will also have a joint live demo between SceniX and Analog Devices, Inc. on real-to-sim-to-real cable manipulation at the ICRA exhibition. This is a small teaser of what we have been building, with more to come soon! If you are at ICRA, please stop by the sessions or the demo booth. Happy to chat about robot learning, simulation, world models, and sim-to-real!show more

Yunzhu Li
11,102 次观看 • 3 个月前
Train a TensorFlow object detection model – then deploy... it on a robot 🤖 Iulia Feroli (Iulia Feroli) shows how to turn a notebook into a real-time object detection app. This tutorial works for any project – though we demonstrate the deployment on (and assisted by!) the #ReachyMini, an open-source robot from Pollen Robotics. Built with PyCharm + Claude Code. 👉 Watch it in action:show more

PyCharm, a JetBrains IDE
34,908 次观看 • 4 个月前
Most imitation learning policies break when the camera moves... or the robot changes. NOT THIS ONE 👇 [📍 Bookmark for later ] A new 3D scene representation encoder, tackles this by enabling zero-shot generalization to unseen embodiments and viewpoints… And it works with any IL algorithm. The trick? •Use a 2D foundation model to extract semantic features •Lift them into 3D space for localization (not semantics) •Condition the IL policy on this spatially grounded vector Across 93 simulated and 6 real tasks, Adapt3R: ✅ Maintains IL performance on LIBERO & MimicGen benchmarks ✅ Outperforms DP3 and 3D Diffuser Actor in most settings ✅ Holds >80% success on LIBERO even with large camera rotations Thanks for sharing this, Animesh Garg & Albert Wilcox! 📍Paper: Website: Code:show more

Ilir Aliu
12,178 次观看 • 1 年前
🤖 Another zero-shot reward model is now in LeRobot:... ROBOMETER. A general-purpose, zero-shot video-language reward model from University of South Carolina, UT Dallas, Massachusetts Institute of Technology (MIT), University of Washington, Ai2, and NVIDIA that predicts frame-level task progress. Trained on 1M+ trajectories from 21 robot embodiments, generalizes zero-shot to unseen tasks, scenes, and robots. 2.4–4.5x better downstream success rates across online RL, offline RL, data filtering, failure detection, and data retrieval for IL. Project: Paper:show more

LeRobot
32,625 次观看 • 3 个月前
Robot Utility Models (RUMs) enable basic tasks – door... opening, drawer opening, object reorientation, etc. – at ~90% accuracy without ANY finetuning (i.e. zero-shot) in unseen new environments. Fully open source!!! models, data, code & hw. We think this is super exciting, why?👇 1. Unlocks many practical home utility tasks that often involve these basic tasks as part of an action chain. “Go get me a fork” involves opening the kitchen door and then opening the cutlery drawer. 2. This works well **zero-shot in unseen and new** environments, which is practically a huge deal. Turn the robot on, and get going. 3. The recipe for building a new model is fairly generic, and we think with a bit more refinement this can be a general recipe to build many more Utility models. More details and access 👇show more

Mahi Shafiullah 🏠🤖
89,535 次观看 • 2 年前
30 minutes of video. Robot learns the task. Open-source,... end-to-end. An open-source framework for training robot policies from only 30 minutes of human egocentric videos captured via Meta Aria glasses: Achieving zero-shot transfer to robots without any robot data collection. The method relies on Interaction-Centric Tokens that encode hand-object spatial relationships invariant to embodiment and viewpoint, supplemented by auxiliary objectives like object motion prediction and latent consistency to extract richer supervision signals from the same data. HumanEgo demonstrates strong cross-embodiment, cross-environment performance on bimanual tasks, outperforming baselines like ACT and teleop data while being trainable on a single RTX 4090 GPU. Thanks for sharing, Zhi (Leo) Wang. 📌 Website: Paper: Code: Video: ——- Weekly robotics and AI insights. Subscribe free:show more

Ilir Aliu
17,077 次观看 • 3 个月前
It's been incredible to see neural networks working so... well on our humanoid robots Humanoids are crazy complex - an individual motor can rotate 360 degrees and you have 40+ joints. If you do the math, that means more possible robot states than atoms in the universe Figure has our own AI model called Helix that we've designed in-house. A single Helix neural network now outputs both manipulation and navigation, end-to-end from language and pixel input Every leap in machine learning has come from massive, diverse datasets. At Figure, we’re currently building the largest pretraining dataset for humanoids in history - excited to see what this unlocksshow more

Brett Adcock
93,986 次观看 • 11 个月前
World modeling and imitation learning have largely been considered... two disparate worlds. In our recent work, Unified World Models, just accepted to #RSS2025, Chuning Zhu provides a dead-simple unifying solution: just train a joint diffusion model over actions and future states, but with *decoupled* diffusion time steps across these modalities. Manipulating these decoupled time steps then allows for marginalization or conditioning on actions or states; a single model can serve as a policy, forward dynamics model, video prediction model, or inverse dynamics model by simply setting diffusion timesteps carefully. The resulting model can leverage video datasets along with robot training data much more effectively, and shows improved robustness, generalization, and flexibility. This is exciting because it is frustratingly simple, scalable, and shows strong improvement on real-world robotics problems. Please refer to Chuning Zhu 's excellent thread for more details! More details/code can be found on our website and in the paper -show more

Abhishek Gupta
11,430 次观看 • 1 年前
The next manipulation tool may not live on your... screen. It may stand in front of you look into your eyes and convince you that it understands. Would you rather face a humanoid robot strong enough to hurt you or one designed well enough to make you trust it? This robot copies blinking, eye contact, head movement and facial expressions to create the illusion of human presence. That may look impressive but it also opens a darker question. When a machine can look concerned appear friendly and imitate emotion people may start trusting signals that contain no real feeling no empathy and no moral responsibility. So which is more dangerous a robot with physical power or a robot that can manufacture trust? #HumanoidRobot #Robotics #AIshow more

Techniahqrobot | humanoid robots
11,650 次观看 • 1 个月前
Pi0 vs. ACT with BBox conditioning 🟦 Not many... know you can push ACT to *almost-Pi0* generalisation by conditioning on bounding boxes (BBoxes). How is the training data collected? • Generate BBoxes for all pick-and-place objects in the scene.(I used Gemini) • Pick-and-place targets are selected randomly. • Add the BBox coordinates to the robot’s state. • Overlay the BBoxes in the visualisation so you know what to grab and where to drop. During inference: • Generate BBoxes for every object again. • Click the object you want to pick and its target spot; those BBoxes get added to the robot state. • Let the robot do the work for you 😃 Setup: - Trained ACT for 100k steps and fine-tuned Pi0 for only 20k. - Training data is 60 episodes and had *only* LEGO bricks. - Using single front camera (Laptop in this case) Got the idea from xun in LeRobot discord. Here’s ACT vs Pi0 on a toy car that isn’t in the dataset. 1/3show more

Shreyas Gite
34,863 次观看 • 1 年前
Some random thoughts reading through the new RobbyAnt VLA... paper: - 20,000 hours of data across 9 robots!! - damn, Chinese companies are going to trivially outscale the American ones on real robot data - It's really cool they train on depth; it means they can handle transparent objects really well for example - You can never tell how good these models are without trying them, since everyone trains on different robots, but from the results they show it does seem to clean up - cross embodiment scaling laws are really cool - if you have lots of robots do you need human video??show more

Chris Paxton
19,921 次观看 • 7 个月前
Disappointed with your ICLR paper being rejected? Ten years... ago today, Sergey and I finished training some of the first end-to-end neutral nets for robot control 🤖 We submitted the paper to RSS on January 23, 2015. It was rejected for being "incremental" and "unlikely to have much impact" Our resubmission to NeurIPS was also rejected It now has >4,000 citations (and more importantly, end-to-end training is widely accepted!) It's also cool to think about what's changed and what's the same -- - The network was 92k parameters and trained on ~15 minutes of data - The code was a combination of matlab, caffe, ROS, a custom CUDA kernel for speed, and a low-level 20 Hz controller in C++, all talking to each other. ROS+matlab was as bad as it sounds. - We pre-trained the encoder and did inference off-board on a workstation with a larger GPU. - We were paranoid about varying lighting messing up the network, so we did all the experiments after sunset (so long nights running experiments on the robot past 3 am) Now, we have manipulation policies that are far more dextrous, far more generalizable, and maybe on the cusp of breaking into the real world. :) (the paper:show more

Chelsea Finn
169,288 次观看 • 1 年前