Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Diffusion and flow matching-based robot planners are slow and generate noisy and jerky trajectories. Delighted to share our ICRA 2026 paper, which leverages IMLE to improve planning frequency 19-fold from 4.3 Hz to 83 Hz and reduces jerk by 38% relative to flow matching. Joint work w/ Grayson Lee,...

102,877 Aufrufe • vor 4 Monaten •via X (Twitter)

16 Kommentare

Profilbild von Ke Li 🍁 @ ECCV 2026
Ke Li 🍁 @ ECCV 2026vor 4 Monaten

Generative models, such as diffusion models and flow matching, are widely used for robot planning. The goal is to model the probability distribution over plausible future trajectories conditioned on the initial state. (2/7)

Profilbild von Ke Li 🍁 @ ECCV 2026
Ke Li 🍁 @ ECCV 2026vor 4 Monaten

Our method uses Implicit Maximum Likelihood Estimation (IMLE) to train a one-step generative model on trajectories of varying quality. IMLE works by having each trajectory in the training dataset pull the closest generated trajectory towards it. (3/7)

Profilbild von Ke Li 🍁 @ ECCV 2026
Ke Li 🍁 @ ECCV 2026vor 4 Monaten

Vanilla IMLE treats all trajectories in the training dataset equally despite their varying quality. Since our goal is to generate trajectories of high quality, we propose reward-weighted IMLE, which emphasizes trajectories with high reward more than those with low reward. (4/7)

Profilbild von Ke Li 🍁 @ ECCV 2026
Ke Li 🍁 @ ECCV 2026vor 4 Monaten

We demonstrate on offline reinforcement learning tasks that our IMLE model runs in real time on both the CPU and GPU and achieves an order-of-magnitude speedup relative to diffusion models. (5/7)

Profilbild von Ke Li 🍁 @ ECCV 2026
Ke Li 🍁 @ ECCV 2026vor 4 Monaten

We deploy our planner onboard a mobile robot and demonstrate real-time navigation in dynamic multi-agent environments. (6/7)

Profilbild von Ke Li 🍁 @ ECCV 2026
Ke Li 🍁 @ ECCV 2026vor 4 Monaten

For more details, see: - Project website: - Paper: - Code: If you are at ICRA, come check out Poster 301 on Tuesday afternoon! (7/7)

Profilbild von Mr V
Mr Vvor 4 Monaten

Bruh diffusion models are really bad at generalizing in real-time , learnt a lot from your paper and I think I will find myself always going back to this. Who would have thought that the denoising process is the root cause. My initial thoughts were that without the de-noisers generalizations wouldn’t work or be optimal. This is great works and it is real time.

Profilbild von Emeka
Emekavor 4 Monaten

This is so nice ! 👏👏👏

Profilbild von oriel haim
oriel haimvor 4 Monaten

This is really cool! Are you looking for employees? Not for the salary I just like to play with things like this 🙃

Profilbild von Julley Thai
Julley Thaivor 4 Monaten

Lil bro playing crossing road.

Profilbild von Priyav K Kaneria
Priyav K Kaneriavor 4 Monaten

really cool

Profilbild von Mooncast Productions
Mooncast Productionsvor 4 Monaten

boids demo?

Profilbild von Alper Ahmetoglu
Alper Ahmetogluvor 4 Monaten

Felt nostalgic. 2018, I was trying something like IMLE in the other KL direction, and was trying to find a way to prevent collapse. Then I stumbled upon IMLE. Cool idea! Will check this out.

Profilbild von LSMO
LSMOvor 4 Monaten

Cool

Profilbild von Saleh Aldwais
Saleh Aldwaisvor 4 Monaten

This is good it didn’t stop at all butterfly find a way while moving

Profilbild von Wendy Carlosa
Wendy Carlosavor 4 Monaten

looks like how the military would operate if the soldiers were cats

Ähnliche Videos

Meta just announced FlowVid Taming Imperfect Optical Flows for Consistent Video-to-Video Synthesis paper page: Diffusion models have transformed the image-to-image (I2I) synthesis and are now permeating into videos. However, the advancement of video-to-video (V2V) synthesis has been hampered by the challenge of maintaining temporal consistency across video frames. This paper proposes a consistent V2V synthesis framework by jointly leveraging spatial conditions and temporal optical flow clues within the source video. Contrary to prior methods that strictly adhere to optical flow, our approach harnesses its benefits while handling the imperfection in flow estimation. We encode the optical flow via warping from the first frame and serve it as a supplementary reference in the diffusion model. This enables our model for video synthesis by editing the first frame with any prevalent I2I models and then propagating edits to successive frames. Our V2V model, FlowVid, demonstrates remarkable properties: (1) Flexibility: FlowVid works seamlessly with existing I2I models, facilitating various modifications, including stylization, object swaps, and local edits. (2) Efficiency: Generation of a 4-second video with 30 FPS and 512x512 resolution takes only 1.5 minutes, which is 3.1x, 7.2x, and 10.5x faster than CoDeF, Rerender, and TokenFlow, respectively. (3) High-quality: In user studies, our FlowVid is preferred 45.7% of the time, outperforming CoDeF (3.5%), Rerender (10.2%), and TokenFlow (40.4%).

AK

123,729 Aufrufe • vor 2 Jahren

Excited to announce GR00T N1, the world’s first open foundation model for humanoid robots! We are on a mission to democratize Physical AI. The power of general robot brain, in the palm of your hand - with only 2B parameters, N1 learns from the most diverse physical action dataset ever compiled and punches above its weight: - Real humanoid teleoperation data. - Large-scale simulation data: we are open-sourcing 300K+ trajectories! - Neural trajectories: we apply SOTA video generation models to “hallucinate” new synthetic data that features accurate physics in pixels. Using Jensen’s words, “systematically infinite data”! - Latent actions: we develop novel algorithms to extract action tokens from in-the-wild human videos and neural generated videos. GR00T N1 is a single end-to-end neural net, from photons to actions: - Vision-Language Model (System 2) that interprets the physical world through vision and language instructions, enabling robots to reason about their environment and instructions, and plan the right actions. - Diffusion Transformer (System 1) that “renders” smooth and precise motor actions at 120 Hz, executing the latent plan made by System 2. We deploy N1 on GR1 robot, 1X Neo robot, and a large collection of simulation benchmarks. N1 achieves up to +30% boost in diverse manipulation tasks for household and industrial settings. While humanoid robots are the main focus of N1, our model also supports cross-embodiment. We finetune it to work on the $110 HuggingFace LeRobot SO100 robot arm! Open robot brain runs on open hardware. Sounds just right. Let’s solve robotics, together, one token at a time. Links to our Whitepaper, Github repo, HuggingFace model, and open dataset page in the thread: 🧵

Jim Fan

467,596 Aufrufe • vor 1 Jahr

New model: your robot can now pack your suitcase 🧳 Xiaomi has released a new robot foundation model. Called Xiaomi-Robotics-1, it is designed to have a robot pick things up and move them around. But first, DEFINITIONS: - Mixture-of-Transformers (MoT): An architecture where separate transformer "experts" (e.g., one for vision-language, one for actions) share a single attention stream, so each modality gets specialized parameters without losing joint reasoning. - Vision-language model (VLM): A model that jointly understands images and text. - Diffusion transformer: A transformer trained to turn noise into structured outputs by iterative denoising, here generating robot actions rather than images. - Action chunks: Short sequences of future actions (e.g., the next ~50 motor commands) predicted in one shot instead of one step at a time. - Flow matching: A faster version of diffusion. The model learns a straight-line velocity field from noise to the target action, so it needs only a few integration steps instead of many denoising ones. Its peculiarity comes from its two stage training: 1. 100,000 hours of video shot through a UMI rig: a handheld 3D-printed gripper with a camera, worn by humans doing ordinary tasks in homes, shops, factories and offices. 2. Adapt to actual robot bodies with ~10,000 hours of real-robot data. It replaces the standard approach of teleoperating a real robot for every hour of training data. Its architecture is a Mixture-of-Transformers pairing a pre-trained Qwen3-VL vision-language model with a diffusion transformer that emits action chunks via flow matching, released in 2.6B, 5.1B and 10.5B parameter variants. However, if you read the entire paper ("Scaling VLA Models with over 100K Hours"), you realize that all of the scaling experiments on 20k hours. Therefore the headline "out-of-the-box success climbing 26% → 75% as pre-training data grows" tops out at 100% of 20k hours! What the full corpus does to that curve is never shown -> and this where things would become interesting! Xiaomi's own conclusion is that model size has stopped mattering and data is the binding constraint. The performance gap among different model sizes are less pronounced than those observed across different data scales. This result suggests that model capacity at the billions-parameter scale may already be sufficient to capture the current dataset's distribution. Which further asks the same question: why not use the 100k video hours? Anyway, I would definitely love to have a couple robots at home that can cooperate to pack my suitcase with items relevant to my next destination:

Léo

15,777 Aufrufe • vor 2 Monaten

🔥 JUST IN: Open-source robotics dataset from 100% real-world scenarios! 🤯 Chinese robotics company AGIBOT just released AGIBOT WORLD 2026, an open-source dataset systematically covering key embodied AI research directions. Built entirely from real-world environments: commercial spaces, and homes. Collected using AGIBOT G2 robots in free-form collection mode, providing structured, accurately annotated, high-quality data. Digital twin technology creates 1:1 scale replicas in simulation matching the real environments. Both real-world and simulation data are open-sourced. The AGIBOT G2 platform collects multiple data types simultaneously: RGB(D) cameras, tactile sensors, force sensors, LiDAR, IMU, and full-body joint states. Whole-body control coordinates arms, waist, and hands for complex tasks. First-person teleoperation lets operators control the robot from its perspective. The tasks covered are fine-grained manipulation, ultra-long-horizon tasks, spatial navigation, dual-arm coordination, and multi-agent/human-robot collaboration. The dataset includes error-recovery trajectories with annotations. Most datasets only show successful demonstrations. AGIBOT includes failures and how the robot recovers, teaching models how to handle mistakes. After collection, data is tested through policy training and real-robot deployment to ensure quality. Then processed through industrial quality control with multiple screening and cleaning rounds. Making it open-source accelerates embodied AI research by giving researchers access to high-quality real-world robot data at scale. 🇨🇳 Learn more here: ~~ ♻️ Join the weekly robotics newsletter, and never miss any news →

Lukas Ziegler

40,583 Aufrufe • vor 6 Monaten

PhD Students – How to automatically identify 90% of the issues in your research paper before you submit it to a journal? This is possible through manual or automated paper review. First, let’s understand the following. 𝐖𝐡𝐚𝐭 𝐢𝐬 𝐚 𝐩𝐚𝐩𝐞𝐫 𝐫𝐞𝐯𝐢𝐞𝐰? Paper review is a process in which subject matter experts evaluate your paper based on the following criteria: 1. Significance – Is this research important? 2. Novelty – Is this research new? 3. Methodology – Is this research carried out in the correct way? 4. Verifiability – Can other researchers verify this research? 5. Presentation – Is the research presented in the right way? 𝐖𝐡𝐲 𝐭𝐨 𝐡𝐚𝐯𝐞 𝐲𝐨𝐮𝐫 𝐩𝐚𝐩𝐞𝐫 𝐫𝐞𝐯𝐢𝐞𝐰𝐞𝐝 𝐛𝐞𝐟𝐨𝐫𝐞 𝐬𝐮𝐛𝐦𝐢𝐬𝐬𝐢𝐨𝐧? ➟ Identify the critical issues in your paper ➟ Fix those issues to increase the chances of your paper acceptance 𝐇𝐨𝐰 𝐭𝐨 𝐚𝐮𝐭𝐨𝐦𝐚𝐭𝐞 “𝐬𝐞𝐥𝐟-𝐫𝐞𝐯𝐢𝐞𝐰” 𝐨𝐟 𝐲𝐨𝐮𝐫 𝐩𝐚𝐩𝐞𝐫? Paperpal just launched an amazing feature – AI Review. Using this feature, you can get instant self-feedback. This feature will help you in the following ways. ➝ Check for gaps in your logic ➝ Get feedback on the structure and flow of your writing ➝ Review your research questions ➝ Identify opportunities to strengthen your paper ➝ Increase the chances of your paper acceptance Here is a step-by-step process for using AI Review feature. Step 1: Go to and login. Step 2: Open an existing document or make a new document Step 3: Go to the right-side bar and click on checks | AI Review. Step 4: For this feature to work there should be more than 150 words. Step 5: Copy and paste your paper. Step 6: Now go to the right side and check the prompts Step 7: With these prompts, you will evaluate your paper. Step 8: You will find various prompts e.g., suggest writing feedback, check flow and structure etc. Step 9: You can select a prompt from the existing prompts or write your custom prompt and execute Step 10: Paperpal will generate feedback as per the prompt. Step 11: Read through the feedback and save it for further use. Use other specific prompts for tailored feedback. Step 12: This way you can evaluate various aspects of your paper yourself. This is a very customized and efficient way of automatically reviewing your paper. You can also go one step further to work on the feedback and improve your paper based on suggestions. Please note that AI Review feature does not replace human or expert reviewers in any way. This feature only aims to provide you with quick self-feedback. Try the AI Review feature of Paperpal. Paperpal link:

Faheem Ullah

15,270 Aufrufe • vor 1 Jahr

We are excited to share our work “Event-Aided Sharp Radiance Field Reconstruction for Fast-Flying Drones” published in IEEE Transactions on Robotics IEEE Transactions on Robotics (T-RO), which tackles sharp radiance field reconstruction under agile drone motion, where RGB frames are heavily motion-blurred and pose priors become unreliable! 4 years in the making! Code & dataset released! PDF: Code & Dataset: Full Narrated Video: High-speed flight is essential for time- and battery-constrained missions (e.g., inspection, exploration, search & rescue). However, fast motion corrupts visual data with severe motion blur and introduces drift/noise in visual-inertial odometry, making NeRF-based 3D reconstruction particularly brittle. We propose a unified framework that leverages asynchronous #EventCamera streams together with motion-blurred frames to reconstruct high-fidelity radiance fields from agile drone flights. Our key idea is to embed event-image fusion directly into radiance field optimization while jointly refining a shared, continuous-time camera trajectory initialized from event-based VIO. This enables us to recover sharp radiance fields and accurate trajectories without ground-truth supervision during training. We validate our method on synthetic data and on real sequences captured by a drone flying up to 2 m/s. Despite severe blur and noisy pose priors, our method preserves fine scene details and achieves a performance gain of over 50% on real-world data compared to state-of-the-art methods. Kudos to Rong Zou and Marco Cannici! Marco Cannici Reference: Rong Zou*, Marco Cannici*, Davide Scaramuzza Event-Aided Sharp Radiance Field Reconstruction for Fast-Flying Drones IEEE Transactions on Robotics (T-RO), 2026 NCCR Robotics European Research Council (ERC) AUTOASSESS UZH IfI University of Zurich UZH Science Prophesee SynSense UZH Space Hub

Davide Scaramuzza

12,028 Aufrufe • vor 7 Monaten