Why is action chunking crucial for robot dexterity? 🤖... - We identify a natural tradeoff between temporal consistency and reactivity - New policy decoding technique that is *both* temporally consistent & fully reactive ICLR 2025 paper: A short thread 🧵show more

Chelsea Finn
36,073 просмотров • 1 год назад
How can we more effectively leverage robot data from... different embodiments for skill transfer? Excited to share that our new work, RoVi-Aug, has been accepted to Conference on Robot Learning as an oral paper! WIth RoVi-Aug, you can augment an existing robot dataset into a different robot and different viewpoints. A policy trained on the augmented dataset can zero-shot deploy on the unseen target robot with significantly different camera angles! 🧵👇 🔗 Check out our paper:show more

Chenfeng_X
28,321 просмотров • 2 лет назад
We’ve only had access to GAIA-1 for a few... weeks and are discovering new capabilities every day. The results are phenomenal! GAIA isn't just a generative video model, it is a world model, ie. controllable by video, text & action prompts Why is this huge for self driving? Thread:show more

Alex Kendall
185,472 просмотров • 3 лет назад
Graham Stewart says the SNP's policy for Holyrood 2026... is that the only mandate for an Indyref is an SNP majority. FALSE! The SNP's policy is that a majority of MSPs and a majority for the SNP are equal mandates but that precedent shows the UK Govt cannot ignore the latter. We knew BBC Scotland would misrepresent the policy sooner or later. It's why we recorded Angus Robertson being interviewed last month where he made it very clear that both scenarios were equal in terms of mandate.show more

MSM Monitor
15,569 просмотров • 10 месяцев назад
After owning New Mexico's only geothermal power plant for... 1 yr (as of this week 🥳), we (Zanskar) have fully repowered the plant back to it's nameplate capacity after drilling, completing and tying-in a new (monster) production well Below are some follow-on thoughts about how we did this and why this is just the beginning 🧵show more

Joel Edwards
11,427 просмотров • 1 год назад
There's a popular rumor that storyboards completely ruin character... consistency in AI animation. Well, yes and no. 😅 For this video, I generated myself in a bright new look, created a sketch storyboard in GPT Image 2, and animated it with Seedance 2.0. From my experience, if your storyboard is in hyper-realism, the likeness definitely tends to drift. But if you keep it strictly as a sketch? The face stays surprisingly consistent. I’m sharing the exact prompts I used for both the storyboard and the animation in the thread below! 👇show more

Ivanna | AI Art & Prompts
38,317 просмотров • 2 месяцев назад
Robot Utility Models (RUMs) enable basic tasks – door... opening, drawer opening, object reorientation, etc. – at ~90% accuracy without ANY finetuning (i.e. zero-shot) in unseen new environments. Fully open source!!! models, data, code & hw. We think this is super exciting, why?👇 1. Unlocks many practical home utility tasks that often involve these basic tasks as part of an action chain. “Go get me a fork” involves opening the kitchen door and then opening the cutlery drawer. 2. This works well **zero-shot in unseen and new** environments, which is practically a huge deal. Turn the robot on, and get going. 3. The recipe for building a new model is fairly generic, and we think with a bit more refinement this can be a general recipe to build many more Utility models. More details and access 👇show more

Mahi Shafiullah 🏠🤖
89,535 просмотров • 2 лет назад
𝗥𝗼𝗯𝗼𝘁𝘀 𝗱𝗼𝗻’𝘁 𝗻𝗲𝗲𝗱 𝗺𝗼𝗿𝗲 𝗱𝗲𝗺𝗼𝗻𝘀𝘁𝗿𝗮𝘁𝗶𝗼𝗻𝘀. 𝗧𝗵𝗲𝘆 𝗻𝗲𝗲𝗱 𝘁𝗼 𝗹𝗲𝗮𝗿𝗻... 𝗳𝗿𝗼𝗺 𝗳𝗮𝗶𝗹𝘂𝗿𝗲 — 𝗮𝗳𝘁𝗲𝗿 𝘄𝗮𝘁𝗰𝗵𝗶𝗻𝗴 𝗵𝘂𝗺𝗮𝗻𝘀. Most robot learning systems assume failure is the end of learning. In our new work, we study whether robots can improve after deployment by learning from their own failures, without any human intervention, teleoperation, or corrective labels. The key idea is simple: human videos contain structure about how the world works. We use them to learn cross-embodiment representations of action, dynamics, and value, enabling a shared predictive space between human behavior and robot experience. This allows a new learning loop: 👉 pretrain on human videos 👉 deploy robot policy 👉 observe failures 👉 reinterpret failures using human priors 👉 improve autonomously We evaluate this across 7 real-world manipulation tasks, showing: 📈 40% → 81% success rate 🏆 Strong improvements over π0.6 RECAP and RISE ✔️ Zero human intervention during post-deployment improvement 🧬 Generalizes across robot embodiments and policy backbones A key finding is that explicit failure repair significantly outperforms failure reweighting, yielding substantially larger gains under identical data conditions (+25 pts vs +5 pts on the same π0.5 base policy). Overall, the results suggest a shift in how we think about robot learning: Human videos are not only for pretraining policies. They can provide the structure needed for continual self-improvement after deployment. 📄 Paper: 🌐 Project: I am grateful for working with the fantastic leads Hanzhi Chen and Anran Zhang, and our collaborators Simon Schaefer, Kejia Chen, Shi Chen, Daniel Cremers. Special thanks to Stefan Leutenegger for co-advising this project with me. ETH Zürich TU München Microsoft Check out Hanzhi's 🧵 for more detailsshow more

Oier Mees
12,379 просмотров • 2 месяцев назад
Disappointed with your ICLR paper being rejected? Ten years... ago today, Sergey and I finished training some of the first end-to-end neutral nets for robot control 🤖 We submitted the paper to RSS on January 23, 2015. It was rejected for being "incremental" and "unlikely to have much impact" Our resubmission to NeurIPS was also rejected It now has >4,000 citations (and more importantly, end-to-end training is widely accepted!) It's also cool to think about what's changed and what's the same -- - The network was 92k parameters and trained on ~15 minutes of data - The code was a combination of matlab, caffe, ROS, a custom CUDA kernel for speed, and a low-level 20 Hz controller in C++, all talking to each other. ROS+matlab was as bad as it sounds. - We pre-trained the encoder and did inference off-board on a workstation with a larger GPU. - We were paranoid about varying lighting messing up the network, so we did all the experiments after sunset (so long nights running experiments on the robot past 3 am) Now, we have manipulation policies that are far more dextrous, far more generalizable, and maybe on the cusp of breaking into the real world. :) (the paper:show more

Chelsea Finn
169,288 просмотров • 1 год назад
Hey #NeuraxonMini is literally out! , we manage to... "transplant" a Neuraxon 2 bioinspired #AI brain to a physical robot the #SpheroMini moving from our last Scientific Paper (link bellow) by David Vivancos - e/acc & Jose Sánchez for Qubic #OpenScience hybridized with #Aigarth to the real World. First you need a Sphero Education Mini robot about 50$ Then you can try the first cool demos at Hugging Face: 1.- Neuraxon2MiniControl to drive the sphero robot 2.- Neuraxon2MiniWrite to write letters or words with physical moves of the sphero robot using Neuraxon Video Tutorials on youtube later today. Why this matters? Remember we are not building "dead" LLMs we are building #AliveAIs and for that we need to explore how it behaves in reality, from how it learns to how it fails, and what better way that in the emerging field of #robotics , time will tell if your next #HumanoidRobot have a #Neuraxon brain... Read the Paper: Explore the Neuraxon code here: Are you ready for #TrueAI ?show more

David Vivancos - e/acc
29,293 просмотров • 6 месяцев назад
Most robots still need markers, checkerboards, or long calibration... rituals just to know where their arms are. Now it works from raw images in seconds. roboreg is a markerless multi arm localization toolkit that plugs into ROS 2 and RViz. No special hardware. No custom setup. You toggle between robot descriptions and the system figures out the rest. The idea is simple: ✅ Hand eye calibration from plain RGB or RGB D images ✅ Only three robot poses needed for millimeter accuracy ✅ Works with any ROS 2 compatible robot and camera ✅ Fully open source under Apache 2.0 It is powered by Hydra, a new marker free ICP variant that converges far more reliably than classical baselines and runs in under a second. If you want to try it: roboreg: ROS 2 roboreg: Hydra paper: pip install roboreg More details and discussion on Open Robotics Discourse:show more

Ilir Aliu
18,406 просмотров • 9 месяцев назад
𝗗𝗼𝗻'𝘁 𝗳𝗶𝗻𝗲-𝘁𝘂𝗻𝗲 𝗿𝗼𝗯𝗼𝘁 𝗳𝗼𝘂𝗻𝗱𝗮𝘁𝗶𝗼𝗻 𝗺𝗼𝗱𝗲𝗹𝘀. 𝗦𝘁𝗲𝗲𝗿 𝘁𝗵𝗲𝗺 𝘄𝗶𝘁𝗵 𝗵𝘂𝗺𝗮𝗻... 𝗰𝗼𝗿𝗿𝗲𝗰𝘁𝗶𝗼𝗻𝘀 𝗶𝗻𝘀𝘁𝗲𝗮𝗱, 𝘄𝗶𝘁𝗵𝗼𝘂𝘁 𝗰𝗵𝗮𝗻𝗴𝗶𝗻𝗴 𝘁𝗵𝗲 𝗯𝗮𝘀𝗲 𝗽𝗼𝗹𝗶𝗰𝘆 Modern VLAs and world-action models can perform impressive manipulation skills, but adapting them reliably to new robots and tasks remains challenging. A natural solution is DAgger-style online imitation learning: deploy the robot, collect human corrections, and update the policy. Yet foundation models are fragile in the low-data regime, fine-tuning on a handful of interventions can improve one behavior while degrading others. Online post-training or reinforcement learning can require costly data collection and exploration, making real-world learning expensive and potentially unsafe. In our new paper, 𝗙𝗹𝗼𝘄𝗗𝗔𝗴𝗴𝗲𝗿, we take a different approach: 𝗜𝗻𝘀𝘁𝗲𝗮𝗱 𝗼𝗳 𝗰𝗵𝗮𝗻𝗴𝗶𝗻𝗴 𝘁𝗵𝗲 𝗳𝗼𝘂𝗻𝗱𝗮𝘁𝗶𝗼𝗻 𝗺𝗼𝗱𝗲𝗹, 𝘄𝗲 𝗹𝗲𝗮𝗿𝗻 𝗵𝗼𝘄 𝘁𝗼 𝘀𝘁𝗲𝗲𝗿 𝗶𝘁 𝗳𝗿𝗼𝗺 𝗵𝘂𝗺𝗮𝗻 𝗰𝗼𝗿𝗿𝗲𝗰𝘁𝗶𝗼𝗻𝘀. The key idea is 𝗮𝗰𝘁𝗶𝗼𝗻 𝗶𝗻𝘃𝗲𝗿𝘀𝗶𝗼𝗻: we map human corrective actions back into the latent noise space of the frozen generative policy. These latent targets train a lightweight controller that adapts the robot while preserving the original model's capabilities. Across simulation and real robots, FlowDAgger: 📈 Learns from only 5–20 human intervention episodes 🏆 Outperforms supervised fine-tuning and latent-space reinforcement learning 🤖 Works across VLAs, diffusion policies, and world-action models ✔️ Provides reliable improvements without modifying the pretrained policy We believe this offers a practical path toward making robot foundation models improve during deployment, learning from the way humans naturally teach: through corrections. 📄 Paper: 🌐 Project: 💻 Code: This project was led by my amazing colleague Michael Murray with help from Daphne Chen, Simran Bagaria, Dean Fortier, Tess Hellebrekers, Harshavardhan Reddy Gajarla, Galen Mullins and Andrey Kolobov at Microsoft Research and Maya Cakmak at University of Washingtonshow more

Oier Mees
13,032 просмотров • 1 месяц назад
this guy 3D printed and vibe coded a tiny... Claude robot for his desk it's called "Clawd Mochi." runs on an ESP32 chip with a tiny display that shows animated expressions. > hosts its own WiFi hotspot. zero cloud and zero internet required. fully offline > live-switch between animated faces, a terminal emulator, and a drawing canvas from your browser > total cost: under $8 > takes less than an hour to build 3D print files AND the full build are both open source too this is the greatest thing anyone has built with vibe codingshow more

Om Patel
154,591 просмотров • 5 месяцев назад
A drone that flies, drives, and switches modes in... 0.1 seconds: [Build it yourself: CAD + parts ⬇️] No extra actuators, no deformation, just clever mechanics and full control. DUAWLFIN is a ground-aerial robot with unified actuation: flying like a quadcopter, rolling like a car, and transitioning seamlessly between modes. ✅ Climbs 30° slopes ✅ Hits 2 m/s on wheels with just 15W ✅ Only 3% added energy in flight mode ✅ Mode switch in 0.1s ✅ Fully open-source and 3D-printable Perfect for urban logistics, indoor nav, or just rethinking what drones can be. Paper: Website: Build it yourself: CAD + parts list in the paper 📍 BOOKMARK FOR LATER This is how you merge air and ground without compromise. —- Weekly robotics and AI insights. Subscribe free:show more

Ilir Aliu
69,193 просмотров • 8 месяцев назад
Encouraging to see President Donald J. Trump and America’s... top business leaders making real progress with China. The announcement that China will be purchasing Boeing jets is a major win for American manufacturing, American jobs, and the U.S. economy. I recently visited China, and whether people want to admit it or not, the level of innovation and technological advancement there is impressive. We have to live in reality, and the reality is that China is a major global superpower. Pretending otherwise does not help America compete or lead. When I saw a Chinese robot change its own battery, I realized we have to catch up in some areas. That is why President Trump’s trip is important, not just for trade, but for America’s technological future and competitiveness. America should always put its interests first, but a stable and productive relationship between the world’s two largest economies benefits everyone.show more

Eric Adams
16,344 просмотров • 3 месяцев назад
"Mega Man X5 Improvement Project Addendum" is about to... become "Mega Man X5 Tweaks." Long overdue. The first release aims to make most changes from the Addendum project available individually, including those additional "hex editor" options that some of you might have missed entirely. And everything I've been working on for the past couple of years: 🤖 Armor sprites displayed by Parts (the big one) 🤖 Special Weapons (regular and charged) available for all Armors 🤖 Boss Parameters (Health Bars, Leveling System options, damage and speed values for select moves, and reduced idle times) And a few localization choices that previously were impossible to mix and match: All the English scripts we've gone through: 🤖 Original PS1 version 🤖 Original PS1 + MMXLC2 changes (what we had up to v1.5) 🤖 Retranslation (what we've had since v1.6), now available with either the new or the original font And also: 🤖 Japanese, with all the necessary graphics changes included ...And a bunch of Title Screen options. Why not. The GUI is still a work in progress, obviously, but it's growing daily. I might include a couple of extra things too. The important thing is that every issue under the hood that prevented this from being possible from the start has already been solved. See you soonnnshow more

acediez
34,080 просмотров • 5 месяцев назад
So TeamYouTube has given me a warning and a... strike in 24 hours for two separate videos under the "Harmful and Dangerous Acts" policy and "Harassment and Bullying" policy. I want you to just take a look at what was taken down with the timestamps TeamYouTube has provided A video where i am reacting to a video that is currently on their platform and a video where I'm saying "top 5 gooners" got taken down under this policy with both appeals getting rejected in under an hour TeamYouTube can i get a REAL HUMAN to look at these videos fairly and not a robot. Not one of those links saying "We have previously looked at the video and unfortunately it goes against our community guidelines" the appeal system does not work correctly.show more

Blueryai🌊
10,127 просмотров • 2 месяцев назад
Yesterday, we deployed seven biodegradable natural latex balloons with... measured sulfur dioxide and hydrogen. All reached the stratosphere. Telemetry confirmed 10,525 Cooling Credits reached 20km+, offsetting 10,525 metric tons of carbon dioxide warming for a year. This is the equivalent of 478,409 mature trees that last for a year done in 2 hours. Here is a thread on all seven launches with details: 1st of 7 Balloons Launch Date: July 23, 2025 Balloon Weight: 1,515 grams SO2 Payload: 1,385 grams Injection Altitude: 23,755 meters Delivery: Kaymont HAB-TX-1500 weather balloonshow more

Make Sunsets
79,543 просмотров • 1 год назад
Depth Any Video with Scalable Synthetic Data AI physicists... and chemists continue to make strides in depth estimation from video. Check out this new paper featuring some impressive examples. See the thread for more details (unfortunately no code yet). Abstract: Video depth estimation has long been hindered by the scarcity of consistent and scalable ground truth data, leading to inconsistent and unreliable results. In this paper, we introduce Depth Any Video, a model that tackles the challenge through two key innovations. First, we develop a scalable synthetic data pipeline, capturing real-time video depth data from diverse game environments, yielding 40,000 video clips of 5-second duration, each with precise depth annotations. Second, we leverage the powerful priors of generative video diffusion models to handle real-world videos effectively, integrating advanced techniques such as rotary position encoding and flow matching to further enhance flexibility and efficiency. Unlike previous models, which are limited to fixed-length video sequences, our approach introduces a novel mixed-duration training strategy that handles videos of varying lengths and performs robustly across different frame rates 0 - even on single frames. At inference, we propose a depth interpolation method that enables our model to infer high-resolution video depth across sequences of up to 150 frames. Our model outperforms all previous generative depth models in terms of spatial accuracy and temporal consistency.show more

MrNeRF
27,428 просмотров • 1 год назад
Figure 03 just finished an 8-hour work livestream, imperfect,... but already good enough to replace a lot of repetitive warehouse labor. 🤖 Brett Adcock put a team of F.03 robots on a factory-style package sorting task for a full shift. The job was simple and brutal: detect the barcode, pick the package, flip it label-side down, place it on the conveyor, repeat. Soft poly bags, rigid boxes, moving belts, messy orientations. That is exactly the kind of boring physical work factories pay humans to do all day. Early in the stream, the system handled 230 packages in 10 minutes. That is roughly 2.6 seconds per item — already in human-speed territory for this narrow workflow. The more important part: it was not one robot pretending to work all day. It was a team of Figure 03 robots keeping the line running. When one robot ran low on battery, it left the station and another robot stepped in. That is the real factory signal: not just autonomy, but shift continuity. F.03 is rated for about 5 hours of runtime, so the 8-hour result depends on fleet orchestration, charging, and handoff. That matters more than a single clean demo. The stream was not perfect. There were pauses, hesitations, missed orientations, and small recovery moments. Good. A perfect short clip hides failure. An 8-hour livestream exposes the parts that actually matter: endurance, recovery, throughput, and whether the robot can stay useful after the novelty wears off. Figure says this was fully autonomous on Helix-02, with zero human intervention. For logistics and manufacturing, that is the threshold worth watching. Not “can it do one impressive task?” Can it keep doing the boring task for an entire shift? Figure is not showing a general human replacement yet. But for structured, repetitive factory work, the gap just got much smaller. The timing is also interesting: Figure says BotQ has already delivered 350+ F.03 units and reached a 1 robot/hour production cadence. And F.04 is now in full design lock, with parts starting to ship. The next test is obvious. 8 hours was the proof of endurance. 24/7 is the proof of labor economics.show more

RoboHub🤖
16,818 просмотров • 3 месяцев назад
As we enter New Year 2025, it's reported that... lighting struck the Capitol and the Washington Monument in DC, while in NYC, the Empire State building and the new World Trade Center were also struck. FOUR lighting strikes in one night! Coincidence? Perhaps. But consider that New York is representative of the business and financial world, and Washington DC is the seat of the US government. Is God telling us a big shake-up may be coming soon in both the business and political realms? Interestingly, president-elect Trump is from both worlds. He is the former New York real estate business tycoon and reality TV star turned president who has promised to reform business as usual in Washington as he returns to the presidency. The lighting may signify this kind of earthly shakeup, or nothing at all. Or perhaps it is the omen of something far more consequential, a spiritual shakeup God is planning for America and elsewhere, to arouse many from slumber. We pay much attention to the state of the economy and politics, but are we ready if Christ should return in 2025 as Judge and King, to reckon with us all? Scripture tells us, "Set your minds on things that are above, not on things that are on earth." Colossians 3:2 Now that's wisdom from above!show more

TrueTruth
51,145 просмотров • 1 год назад