Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Can robots learn without trainingโ“ [๐—œ๐˜'๐˜€ ๐—ผ๐—ฝ๐—ฒ๐—ป ๐˜€๐—ผ๐˜‚๐—ฟ๐—ฐ๐—ฒ๐—ฑ โฌ‡ ] Teaching robots to do complex tasks WITHOUT spending hours training them. Sounds cool, right? That's exactly what DIAL-MPC does! The first training-free method for whole-body torque control using full-order dynamics: โœ… Instantly checks if a robot's moves are right...

71,502 Aufrufe โ€ข vor 1 Jahr โ€ขvia X (Twitter)

0 Kommentare

Keine Kommentare verfรผgbar

Kommentare vom Original-Post werden hier angezeigt

ร„hnliche Videos

This work makes a humanoid robot do simple parkour moves by looking with a depth camera and choosing the right move on the fly. The big deal is that it turns lots of small human moves into long, real-time robot behavior, without hand-coding every transition or retraining for each new course. A humanoid robot is usually good at steady walking, but it often fails when it has to do fast moves like jumping up, vaulting, or rolling, and then keep going to the next obstacle. The hard part is that you cannot easily collect training data for every possible obstacle shape, distance, and mistake, so robots end up learning a few moves that only work in a narrow setup. This work starts from short clips of real human parkour moves, like stepping over, vaulting, climbing, and rolling. It uses motion matching, which is basically a smart โ€œpick the next clip that fits best right nowโ€ search, to stitch those short clips into a long, smooth plan that looks like a human doing a whole course. Then it trains a controller with reinforcement learning (RL), which means the robot learns by trial and error to copy that plan while staying balanced and not falling. After training separate expert controllers for different moves, it compresses them into 1 controller that uses only onboard depth sensing and a simple โ€œgo this fast in this directionโ€ command. In real tests on a Unitree G1 humanoid, it can clear multiple obstacles in a row, adapt when obstacles get moved, and climb a wall up to 1.25m.

Rohan Paul

37,121 Aufrufe โ€ข vor 6 Monaten

The term "continual learning" has become overloaded if you see it as an ML problem. One classic thread is about memorization: regularization-based continual learning methods, such as EWC, MAS, and SI, estimate which parameters mattered for previous tasks and resist changing them too much. One modern thread is about adaptation: test-time training and inference-time learning methods, such as TTT, adapt part of the model on the incoming test stream before making predictions. These are sometimes discussed as separate threads. But in modern scalable architectures, I think they are better seen as complementary constraints: a model that learns quickly at test time also benefits from a mechanism for deciding what not to forget. In our #ECCV2026 paper, we study this in large-scale 4D reconstruction: how to build fast spatial memory that can adapt over long observation streams while reducing collapse and forgetting. Instead of using fully plastic test-time updates, we stabilize fast-weight adaptation with an elastic prior that balances adaptation and memory. Key ideas: - Elastic Test-Time Training: Fisher-weighted consolidation for fast-weight updates - EMA anchor weights that provide a moving reference for stability - Chunk-by-chunk inference for long 3D/4D observation streams We show that this scales across large 3D/4D pretraining settings, including both LRM-style and LVSM-style models, and improves reconstruction across benchmarks including Stereo4D, NVIDIA, and DL3DV-140. We release model checkpoints across different design choices: resolution, post-training curriculum, and whether the model uses an explicit 4DGS intermediate representation. - Homepage: - Paper: - Code: - Models: This work is co-led with Xueyang Yu, contributed by Haoyu Zhen Yuncong Yang, and advised by Michigan SLED Lab Chuang Gan.

Martin Ziqiao Ma

33,698 Aufrufe โ€ข vor 2 Monaten

๐——๐—ผ๐—ป'๐˜ ๐—ณ๐—ถ๐—ป๐—ฒ-๐˜๐˜‚๐—ป๐—ฒ ๐—ฟ๐—ผ๐—ฏ๐—ผ๐˜ ๐—ณ๐—ผ๐˜‚๐—ป๐—ฑ๐—ฎ๐˜๐—ถ๐—ผ๐—ป ๐—บ๐—ผ๐—ฑ๐—ฒ๐—น๐˜€. ๐—ฆ๐˜๐—ฒ๐—ฒ๐—ฟ ๐˜๐—ต๐—ฒ๐—บ ๐˜„๐—ถ๐˜๐—ต ๐—ต๐˜‚๐—บ๐—ฎ๐—ป ๐—ฐ๐—ผ๐—ฟ๐—ฟ๐—ฒ๐—ฐ๐˜๐—ถ๐—ผ๐—ป๐˜€ ๐—ถ๐—ป๐˜€๐˜๐—ฒ๐—ฎ๐—ฑ, ๐˜„๐—ถ๐˜๐—ต๐—ผ๐˜‚๐˜ ๐—ฐ๐—ต๐—ฎ๐—ป๐—ด๐—ถ๐—ป๐—ด ๐˜๐—ต๐—ฒ ๐—ฏ๐—ฎ๐˜€๐—ฒ ๐—ฝ๐—ผ๐—น๐—ถ๐—ฐ๐˜† Modern VLAs and world-action models can perform impressive manipulation skills, but adapting them reliably to new robots and tasks remains challenging. A natural solution is DAgger-style online imitation learning: deploy the robot, collect human corrections, and update the policy. Yet foundation models are fragile in the low-data regime, fine-tuning on a handful of interventions can improve one behavior while degrading others. Online post-training or reinforcement learning can require costly data collection and exploration, making real-world learning expensive and potentially unsafe. In our new paper, ๐—™๐—น๐—ผ๐˜„๐——๐—”๐—ด๐—ด๐—ฒ๐—ฟ, we take a different approach: ๐—œ๐—ป๐˜€๐˜๐—ฒ๐—ฎ๐—ฑ ๐—ผ๐—ณ ๐—ฐ๐—ต๐—ฎ๐—ป๐—ด๐—ถ๐—ป๐—ด ๐˜๐—ต๐—ฒ ๐—ณ๐—ผ๐˜‚๐—ป๐—ฑ๐—ฎ๐˜๐—ถ๐—ผ๐—ป ๐—บ๐—ผ๐—ฑ๐—ฒ๐—น, ๐˜„๐—ฒ ๐—น๐—ฒ๐—ฎ๐—ฟ๐—ป ๐—ต๐—ผ๐˜„ ๐˜๐—ผ ๐˜€๐˜๐—ฒ๐—ฒ๐—ฟ ๐—ถ๐˜ ๐—ณ๐—ฟ๐—ผ๐—บ ๐—ต๐˜‚๐—บ๐—ฎ๐—ป ๐—ฐ๐—ผ๐—ฟ๐—ฟ๐—ฒ๐—ฐ๐˜๐—ถ๐—ผ๐—ป๐˜€. The key idea is ๐—ฎ๐—ฐ๐˜๐—ถ๐—ผ๐—ป ๐—ถ๐—ป๐˜ƒ๐—ฒ๐—ฟ๐˜€๐—ถ๐—ผ๐—ป: we map human corrective actions back into the latent noise space of the frozen generative policy. These latent targets train a lightweight controller that adapts the robot while preserving the original model's capabilities. Across simulation and real robots, FlowDAgger: ๐Ÿ“ˆ Learns from only 5โ€“20 human intervention episodes ๐Ÿ† Outperforms supervised fine-tuning and latent-space reinforcement learning ๐Ÿค– Works across VLAs, diffusion policies, and world-action models โœ”๏ธ Provides reliable improvements without modifying the pretrained policy We believe this offers a practical path toward making robot foundation models improve during deployment, learning from the way humans naturally teach: through corrections. ๐Ÿ“„ Paper: ๐ŸŒ Project: ๐Ÿ’ป Code: This project was led by my amazing colleague Michael Murray with help from Daphne Chen, Simran Bagaria, Dean Fortier, Tess Hellebrekers, Harshavardhan Reddy Gajarla, Galen Mullins and Andrey Kolobov at Microsoft Research and Maya Cakmak at University of Washington

Oier Mees

13,032 Aufrufe โ€ข vor 1 Monat

The future of housework just leaked on GitHub and nobody is talking about it. knox byte just open sourced a framework that coordinates swarms of Unitree G1 humanoid robots to clean your entire house on their own. It's called ARGOS. You tell it "clean the bedroom" in plain English and 2+ G1 robots split the room into zones, sweep in parallel, and sync up for the tasks that need four hands like making the bed or moving furniture. The Claude API decomposes your sentence into a task graph. An auction system makes every robot bid on every task based on distance, battery, and current load. The cheapest robot wins. Cooperative jobs go to the cheapest team. Here's what makes this different from every demo video Boston Dynamics keeps teasing: โ†’ 12 cleaning tasks baked in sweeping, mopping, wiping, vacuuming, taking out trash, making the bed, changing sheets, moving furniture, sorting items โ†’ 3 policy architectures running underneath OpenVLA-7B for language tasks, Diffusion Policy for floor coverage, ACT for dexterous bimanual work โ†’ Train it on your own footage record yourself cleaning, run one command, it extracts poses, builds a LeRobot dataset, and LoRA fine-tunes the policy โ†’ PEFA protocol for cooperative work Propose, Execute, Feedback, Adjust. If one robot fails halfway through making the bed, the team replans and retries โ†’ Full MuJoCo simulation so you test policies before pushing them to real hardware โ†’ Silver and cyan terminal dashboard that shows live fleet status, zone maps, task queues, and battery levels in real time The G1 robots talk to each other over CycloneDDS mesh using Unitree's native SDK. No cloud. No middleware. The whole thing runs on a Jetson Orin inside each robot. The wildest part is the training pipeline. Drop cleaning videos into a folder, run argos train ingest, and the framework does the entire pipeline frame extraction, pose estimation, action labeling, HDF5 dataset, fine-tune, evaluate in sim, deploy to robot. One command per stage. Unitree G1s already exist. The framework to make them clean your house just hit GitHub. 52 stars. MIT License. 100% Opensource.

Guri Singh

27,404 Aufrufe โ€ข vor 3 Monaten

Today was my hardest workout before the Javelina 100โ€”an uphill treadmill supercompensation session accumulating 60 minutes of intervals at 10% grade, starting at threshold and ending harder. I think uphill treadmill threshold sessions can be magical for some athletes. Threshold work is classically defined as LT2 or easier, around what you could sustain for 1 hour. In practice, that feels relatively relaxed at first, and it only starts to get harder after you accumulate a substantial amount of volume. The rationale of threshold work is that it improves lactate shuttling, helping mitochondria be more efficient at processing and transporting lactate, preventing fatigue cascades even at harder efforts on other days (or at easier efforts in marathons or ultras). In other words, itโ€™s primarily an aerobic stress. Faster is not better. The real-world obstacles with threshold work are twofold. First, for most of us, itโ€™s pretty slow when done right, or way too hard when done wrong. A study on the training of elite athletes found that long intervals had the lowest correlation with long-term growth, and this conundrum is probably whyโ€”athletes do their long intervals too hard, breaking themselves down without the mechanical or aerobic stimulus to justify it. Second, outdoor threshold work can be an injury risk. If I tried this workout outdoors, it would wreck my calves and high hamstrings for days. The uphill treadmill can help athletes get around these hurdles. Itโ€™s slower by design, putting the emphasis squarely on the aerobic system. That helps athletes develop a much more precise understanding of threshold. But perhaps most significantly, the uphill treadmill reduces impact forces immensely. When I finish one of theseโ€”even a supercompensation sessionโ€”I feel fine the next day, allowing me to absorb way more work (and more specific work to my goals). Particularly with age, I find that running training is about managing the efforts that are high impact to be limited and focused. While most of the uphill treadmill work I do is very controlled, itโ€™s also ok to occasionally dig deeper. Today was about supercompensation. Itโ€™s not called The Pain Cave for nothing ๐Ÿ”ฅ

David Roche

40,681 Aufrufe โ€ข vor 1 Jahr

๐—ฅ๐—ผ๐—ฏ๐—ผ๐˜๐˜€ ๐—ฑ๐—ผ๐—ปโ€™๐˜ ๐—ป๐—ฒ๐—ฒ๐—ฑ ๐—บ๐—ผ๐—ฟ๐—ฒ ๐—ฑ๐—ฒ๐—บ๐—ผ๐—ป๐˜€๐˜๐—ฟ๐—ฎ๐˜๐—ถ๐—ผ๐—ป๐˜€. ๐—ง๐—ต๐—ฒ๐˜† ๐—ป๐—ฒ๐—ฒ๐—ฑ ๐˜๐—ผ ๐—น๐—ฒ๐—ฎ๐—ฟ๐—ป ๐—ณ๐—ฟ๐—ผ๐—บ ๐—ณ๐—ฎ๐—ถ๐—น๐˜‚๐—ฟ๐—ฒ โ€” ๐—ฎ๐—ณ๐˜๐—ฒ๐—ฟ ๐˜„๐—ฎ๐˜๐—ฐ๐—ต๐—ถ๐—ป๐—ด ๐—ต๐˜‚๐—บ๐—ฎ๐—ป๐˜€. Most robot learning systems assume failure is the end of learning. In our new work, we study whether robots can improve after deployment by learning from their own failures, without any human intervention, teleoperation, or corrective labels. The key idea is simple: human videos contain structure about how the world works. We use them to learn cross-embodiment representations of action, dynamics, and value, enabling a shared predictive space between human behavior and robot experience. This allows a new learning loop: ๐Ÿ‘‰ pretrain on human videos ๐Ÿ‘‰ deploy robot policy ๐Ÿ‘‰ observe failures ๐Ÿ‘‰ reinterpret failures using human priors ๐Ÿ‘‰ improve autonomously We evaluate this across 7 real-world manipulation tasks, showing: ๐Ÿ“ˆ 40% โ†’ 81% success rate ๐Ÿ† Strong improvements over ฯ€0.6 RECAP and RISE โœ”๏ธ Zero human intervention during post-deployment improvement ๐Ÿงฌ Generalizes across robot embodiments and policy backbones A key finding is that explicit failure repair significantly outperforms failure reweighting, yielding substantially larger gains under identical data conditions (+25 pts vs +5 pts on the same ฯ€0.5 base policy). Overall, the results suggest a shift in how we think about robot learning: Human videos are not only for pretraining policies. They can provide the structure needed for continual self-improvement after deployment. ๐Ÿ“„ Paper: ๐ŸŒ Project: I am grateful for working with the fantastic leads Hanzhi Chen and Anran Zhang, and our collaborators Simon Schaefer, Kejia Chen, Shi Chen, Daniel Cremers. Special thanks to Stefan Leutenegger for co-advising this project with me. ETH Zรผrich TU Mรผnchen Microsoft Check out Hanzhi's ๐Ÿงต for more details

Oier Mees

12,379 Aufrufe โ€ข vor 1 Monat

You can't 3D reconstruct glass from images... ...WRONG! Thanks for video diffusion, now just about anything is possible! Introducing...Diffusion Knows Transparency (DKT) Transparent and reflective objects usually break robot vision and photogrammetry pipelines because they don't follow the "solid object" rules standard cameras expect. DKT is a new AI model that repurposes the "internal physics engine" found in video generation models to solve this problem. Researchers took a massive video diffusion model (WAN) and fine-tuned it using a custom-built synthetic dataset to turn it into a high-precision depth sensor. To train the AI, they built the first massive synthetic video library of transparent objects, 1.32 million frames of perfectly labeled glass and metal objects in motion. Without ever seeing a "real" labeled video of glass during training, the model (DKT) outperformed all previous specialized systems on real-world benchmarks (ClearPose, DREDS). They created a "lightweight" 1.3B parameter version that runs fast enough (0.17s per frame) to be used on actual robot hardware. Two reasons I find this project important: 1. It further proves that synthetic data will be essential for training the next generation vision models. 2. In real-world robotic tests, using DKT's depth maps nearly doubled the success rate of robot arms trying to pick up objects on tricky reflective or translucent surfaces. At home robots will need to interact with these types of objects on a daily basis. Check out the project page here: Code is LIVE! #Computervision #Robotics #AI

Jonathan Stephens

17,712 Aufrufe โ€ข vor 7 Monaten

Almost 30 years ago my parents moved their young family oversees to help support pastors starting churches. They moved half way across the world and found themselves without a pastor to support. The church in this area was almost nowhere to be found. Rituals replaced relationships and the need for pastors was huge. So Mom and Dad decided to do what they said they wouldnโ€™t do, they started a church. It was just our family at first. Felt kinda silly if Iโ€™m being honest. Just the five of us in our living room. โ€œWhy move across the world for this?โ€ We wondered. Then, as we built relationships, the church began to grow. It was slow at first. It was also hard and lonely. But Dad was bent on building something that would train leaders and would grow without him. He always said our job was to work ourselves out of a job. The goal was to leave the leadership in the hands of the people there as quickly as possible. My parents had to move back to the States rather suddenly a few years ago after my Mom suffered a major stroke during a serious illness. The church was strong and Dad spent years discipline and training leaders. Yet, it was still sudden. We wondered, Would the church continue growing without them or begin to dwindle? Recently I decided to go back home to Northern Chile with my 3 kids to see how the church my parents started in our living room was doing. And look what I found! Itโ€™s a church of well over 500 but thatโ€™s not the impressive part. The church has planted 11 other churches, runs two childrenโ€™s homes, a rehab center, a seminary training hundreds of pastors, and sends people all over the world to work in ministry. Itโ€™s been a rock in the community during earthquakes, tidal waves and civil unrest. I was able to return home and tell Dad, โ€œDad, your church sends their greetingsโ€

Bethany | Commercial Real Estate

27,587 Aufrufe โ€ข vor 1 Jahr

Model-Free Reinforcement Learning (MFRL) has been alluring, especially with supercharged compute with physics on GPU. However, the methods use 0-th order gradients, and are often not the best optimizers. Can we do better than PPO in continuous control for robotics? Turns out yes! ๐Ÿฅณ tl;dr: Faster, better RL than PPO in continuous control ๐Ÿ’ช The answer lies in using more information from the simulation. We are juicing the simulation on GPU as it is, why not use it for gradients as well? This has been a driving question in a series of our works. We first studied this problem in ICLR 2022 paper on Short Horizon Actor Critic Naive gradient based methods are stuck in local minima and have exploding/vanishing gradients. SHAC solved this problem truncated rollouts and model based value estimation, where the model is Differentiable Sim. This boosted sample efficiency and wall-clock time immensely especially in high dimensional systems such as humanoids Yet, given enough compute PPO often caught up. Our follow up paper on on Adaptive Horizon Actor Critic at ICML 2024 discovers the cause and provides a fix. However, we find that even when given ground-truth dynamics, not all gradients are useful due to sample error. 1st-Order Model-Based Reinforcement Learning methods employing differentiable simulation provide gradients with reduced variance but are susceptible to bias in scenarios involving stiff dynamics, such as physical contact. We find that back-propagating through contact and long trajectories drastically reduces gradient accuracy. Using this insight, we propose AHAC to dynamically adapt its roll-out horizon to avoid differentiating through stiff contact. AHAC is a first-order model-based RL algorithm that learns high-dimensional tasks in minutes (wall clock) and outperforms PPO by 40%, even in the limit of data provided to PPO. This work is led by Ignat Georgiev alongside Krishnan Srinivasan, Jie Xu, Eric Heiden and ample assistance from warp team at NVIDIA Robotics (Miles Macklin)

Animesh Garg

52,308 Aufrufe โ€ข vor 2 Jahren