Загрузка видео...

Не удалось загрузить видео

На главную

Why is action chunking crucial for robot dexterity? 🤖 - We identify a natural tradeoff between temporal consistency and reactivity - New policy decoding technique that is *both* temporally consistent & fully reactive ICLR 2025 paper: A short thread 🧵

36,073 просмотров • 1 год назад •via X (Twitter)

Комментарии: 0

Нет доступных комментариев

Здесь появятся комментарии из оригинального поста

Похожие видео

𝗥𝗼𝗯𝗼𝘁𝘀 𝗱𝗼𝗻’𝘁 𝗻𝗲𝗲𝗱 𝗺𝗼𝗿𝗲 𝗱𝗲𝗺𝗼𝗻𝘀𝘁𝗿𝗮𝘁𝗶𝗼𝗻𝘀. 𝗧𝗵𝗲𝘆 𝗻𝗲𝗲𝗱 𝘁𝗼 𝗹𝗲𝗮𝗿𝗻 𝗳𝗿𝗼𝗺 𝗳𝗮𝗶𝗹𝘂𝗿𝗲 — 𝗮𝗳𝘁𝗲𝗿 𝘄𝗮𝘁𝗰𝗵𝗶𝗻𝗴 𝗵𝘂𝗺𝗮𝗻𝘀. Most robot learning systems assume failure is the end of learning. In our new work, we study whether robots can improve after deployment by learning from their own failures, without any human intervention, teleoperation, or corrective labels. The key idea is simple: human videos contain structure about how the world works. We use them to learn cross-embodiment representations of action, dynamics, and value, enabling a shared predictive space between human behavior and robot experience. This allows a new learning loop: 👉 pretrain on human videos 👉 deploy robot policy 👉 observe failures 👉 reinterpret failures using human priors 👉 improve autonomously We evaluate this across 7 real-world manipulation tasks, showing: 📈 40% → 81% success rate 🏆 Strong improvements over π0.6 RECAP and RISE ✔️ Zero human intervention during post-deployment improvement 🧬 Generalizes across robot embodiments and policy backbones A key finding is that explicit failure repair significantly outperforms failure reweighting, yielding substantially larger gains under identical data conditions (+25 pts vs +5 pts on the same π0.5 base policy). Overall, the results suggest a shift in how we think about robot learning: Human videos are not only for pretraining policies. They can provide the structure needed for continual self-improvement after deployment. 📄 Paper: 🌐 Project: I am grateful for working with the fantastic leads Hanzhi Chen and Anran Zhang, and our collaborators Simon Schaefer, Kejia Chen, Shi Chen, Daniel Cremers. Special thanks to Stefan Leutenegger for co-advising this project with me. ETH Zürich TU München Microsoft Check out Hanzhi's 🧵 for more details

Oier Mees

12,277 просмотров • 1 месяц назад

Depth Any Video with Scalable Synthetic Data AI physicists and chemists continue to make strides in depth estimation from video. Check out this new paper featuring some impressive examples. See the thread for more details (unfortunately no code yet). Abstract: Video depth estimation has long been hindered by the scarcity of consistent and scalable ground truth data, leading to inconsistent and unreliable results. In this paper, we introduce Depth Any Video, a model that tackles the challenge through two key innovations. First, we develop a scalable synthetic data pipeline, capturing real-time video depth data from diverse game environments, yielding 40,000 video clips of 5-second duration, each with precise depth annotations. Second, we leverage the powerful priors of generative video diffusion models to handle real-world videos effectively, integrating advanced techniques such as rotary position encoding and flow matching to further enhance flexibility and efficiency. Unlike previous models, which are limited to fixed-length video sequences, our approach introduces a novel mixed-duration training strategy that handles videos of varying lengths and performs robustly across different frame rates 0 - even on single frames. At inference, we propose a depth interpolation method that enables our model to infer high-resolution video depth across sequences of up to 150 frames. Our model outperforms all previous generative depth models in terms of spatial accuracy and temporal consistency.

MrNeRF

27,428 просмотров • 1 год назад

Figure 03 just finished an 8-hour work livestream, imperfect, but already good enough to replace a lot of repetitive warehouse labor. 🤖 Brett Adcock put a team of F.03 robots on a factory-style package sorting task for a full shift. The job was simple and brutal: detect the barcode, pick the package, flip it label-side down, place it on the conveyor, repeat. Soft poly bags, rigid boxes, moving belts, messy orientations. That is exactly the kind of boring physical work factories pay humans to do all day. Early in the stream, the system handled 230 packages in 10 minutes. That is roughly 2.6 seconds per item — already in human-speed territory for this narrow workflow. The more important part: it was not one robot pretending to work all day. It was a team of Figure 03 robots keeping the line running. When one robot ran low on battery, it left the station and another robot stepped in. That is the real factory signal: not just autonomy, but shift continuity. F.03 is rated for about 5 hours of runtime, so the 8-hour result depends on fleet orchestration, charging, and handoff. That matters more than a single clean demo. The stream was not perfect. There were pauses, hesitations, missed orientations, and small recovery moments. Good. A perfect short clip hides failure. An 8-hour livestream exposes the parts that actually matter: endurance, recovery, throughput, and whether the robot can stay useful after the novelty wears off. Figure says this was fully autonomous on Helix-02, with zero human intervention. For logistics and manufacturing, that is the threshold worth watching. Not “can it do one impressive task?” Can it keep doing the boring task for an entire shift? Figure is not showing a general human replacement yet. But for structured, repetitive factory work, the gap just got much smaller. The timing is also interesting: Figure says BotQ has already delivered 350+ F.03 units and reached a 1 robot/hour production cadence. And F.04 is now in full design lock, with parts starting to ship. The next test is obvious. 8 hours was the proof of endurance. 24/7 is the proof of labor economics.

RoboHub🤖

16,818 просмотров • 3 месяцев назад

LEONARDO, also called LEO, was built by researchers at Caltech’s Center for Autonomous Systems and Technologies. Its full name means LEgs ONboARD drOne. The idea is simple but unusual: • Build a small biped robot • Give it drone-style thrust • Use the legs for ground contact • Use the propellers for balance and lift • Combine walking, hopping and flying in one system LEO is basically a hybrid between a walking robot and a flying drone. How it was built: • Two lightweight legs • Three actuated joints in each leg • Four propeller thrusters near the shoulders • A lightweight body • Leg motors for ground movement • Propellers for balance, lift and aerial control • Real-time control software that synchronizes the legs and propellers How it walks: • The legs move the robot forward • The feet touch the ground like a normal biped • The propellers constantly correct balance from above • The robot can stay upright even in unstable situations • The thrust reduces the risk of falling during difficult motions How it flies: • The legs stop being the main locomotion system • The four propellers generate lift • The robot behaves more like a drone • It can take off, fly over obstacles and land back on its legs What makes it different: • It does not walk like a normal humanoid • It does not fly like a normal drone • It blends both systems • The legs handle contact with the ground • The propellers act like fast stabilizers • The control system decides how much help comes from the legs and how much comes from thrust That is why LEO can: • Walk • Hop • Fly over obstacles • Ride a skateboard • Balance on a slackline The key idea is walking with aerial stabilization.

Techniahqrobot | humanoid robots

135,515 просмотров • 1 месяц назад