Загрузка видео...

Не удалось загрузить видео

На главную

A big part of scaling robot learning to solve real-world problems is that we somehow need to get enough diverse, high-quality data to train our robots to perform useful things. GPT and its fellow large language models were bootstrapped and proved out on a massive dataset of real-world language...

20,486 просмотров • 1 год назад •via X (Twitter)

Комментарии: 11

Фото профиля Chris Paxton
Chris Paxton1 год назад

Learning Robotic Manipulation from Simulations A comparison of a few recent works on sim-to-real robot manipulation that I liked

Фото профиля VistaShares
VistaShares1 год назад

From semiconductors to data centers, AIS targets the critical components behind AI's exponential growth. Capture potential returns from this transformative technology sector.

Фото профиля Kyle🤖🚀🦭
Kyle🤖🚀🦭1 год назад

uncontested Isaac superiority 🫡 (very excited for warp too)

Фото профиля ahad
ahad1 год назад

what are you thoughts on a model that you just shovel in alot multimodal data and how much of a ratio you’ll need to get good performance? for example recent work from @physical_int they made an architecture where they can predict web data and predict actions.

Фото профиля Chris Paxton
Chris Paxton1 год назад

Sim to real is definitely not the only way to do this I wrote a previous post that mentioned this; I wrote the two together so they reference the same paper on how much sim data you "need": I think it's worth a deeper look. but my guess is that sim and real data are fulfilling different needs: - sim data is generally very good at capturing robot planning and kinematics - video data is really good at capturing semantics and information about the world - robot data captures everything but isn't diverse enough and is too expensive (although people are trying to change that)

Фото профиля Godwyll Aikins
Godwyll Aikins1 год назад

Excited for real2sim to get easier! Just using your phone + lidar to capture your workspace to port to simulation could be a game-changer for sim2real. Learned simulators could maybe bridge that gap as well

Фото профиля Bahram Banisadr
Bahram Banisadr1 год назад

What do you think about the sim2real transfer strategies & trying to scale up data that way, versus the strategies of training on real, non-robot data, which is readily available? I'm thinking about things like DexMachina learning from human ego-centric demonstration (could put smart glasses on people doing their everyday jobs) or Meta's V-JEPA 2 model that relies on the massive corpus of 'things happening in the world' to build a foundation model that has physical understanding.

Фото профиля Richa Sharma
Richa Sharma1 год назад

@chris_j_paxton I'm building browser-native generalized physics simulators and would love to chat about async agentic acceleration and scaling sim data. The current pipeline is incredibly fragmented.

Фото профиля Chris Paxton
Chris Paxton1 год назад

Yeah interesting!

Фото профиля Martin Matak
Martin Matak1 год назад

yes, your "probably" is correct: we modeled the table geometry as a collision in dextrah (the original one, wasn't part of the rgb extension)

Фото профиля Rawlala
Rawlala1 год назад

the comparison doesn't seem to hold the amount of info u can infer from the universe is literally (uncountably) infinite, and the amount of knowledge we already discover is huge already, hence the necessity for gpt to have access to huge amount of data but for physical task ?..

Похожие видео

Today, we're joined by Nikita Rudin, co-founder and CEO of Flexion to discuss the gap between current robotic capabilities and what’s required to deploy fully autonomous robots in the real world. Nikita explains how reinforcement learning and simulation have driven rapid progress in robot locomotion—and why locomotion is still far from “solved.” We dig into the sim2real gap, and how adding visual inputs introduces noise and significantly complicates sim-to-real transfer. We also explore the debate between end-to-end models and modular approaches, and why separating locomotion, planning, and semantics remains a pragmatic approach today. Nikita also introduces the concept of "real-to-sim", which uses real-world data to refine simulation parameters for higher fidelity training, discusses how reinforcement learning, imitation learning, and teleoperation data are combined to train robust policies for both quadruped and humanoid robots, and introduces Flexion's hierarchical approach that utilizes pre-trained Vision-Language Models (VLMs) for high-level task orchestration with Vision-Language-Action (VLA) models and low-level whole-body trackers. Finally, Nikita shares the behind-the-scenes in humanoid robot demos, his take on reinforcement learning in simulation versus the real world, the nuances of reward tuning, and offers practical advice for researchers and practitioners looking to get started in robotics today. 🗒️ For the full list of resources for this episode, visit the show notes page: 📖 CHAPTERS =============================== 00:00 - Introduction 04:07 - Is robot locomotion solved? 06:04 - Sim-to-real gap 08:58 - Adding semantics to policies 09:42 - Modular vs end-to-end architectures 10:29 - Planner model 12:21 - Adapting RL techniques from quadrupeds to humanoids 15:39 - Behind robot demos 18:09 - Humanoid robots in home environments 22:03 - Training approach 23:56 - VLA models 27:59 - Closing the sim-to-real gap 32:55 - Task orchestration using VLMs 36:38 - Tool use 38:10 - Model hierarchy 43:37 - Simulator versus simulation environment 44:57 - Combining imitation learning and reinforcement learning 46:42 - RL in real world versus RL in simulation 52:58 - Reward tuning and value functions in robotics 56:38 - Predictions 1:00:10 - Humanoids, quadropeds, and wheeled platforms 1:02:45 - Advice, recommended robot kits, and community pla

The TWIML AI Podcast

22,592 просмотров • 7 месяцев назад

Trained a humanoid entirely in a 3D scan of the office. Zero real-world fine-tuning. It just walked in and worked. RL needs hundreds of thousands of attempts, and real robots can't afford to crash. A misjudged gap or a glass door collision breaks hardware and costs hours resetting. So you train in a sim. But sim policies usually train on randomized, untextured geometry; depth is easy to fake. The robot learns structure, not the real world: no materials, no lighting, no idea what anything actually is. RGB cameras carry all of that but training RGB policies in generic fake worlds won’t generalize to the real world. Niantic Spatial 🌎 Scaniverse reconstructs your scan of the real deployment site. One 360° camera walkthrough → photorealistic 3D Gaussian splat at metric scale → collision mesh pulled from the same reconstruction, so vision and physics match exactly. Drops straight into NVIDIA Isaac Sim/Lab, no manual conversion. Flexion simulation-first approach then seamlessly enables the training of RGB-only nav policies inside that reconstruction. With added domain randomization + large image encoders for robustness, this deploys straight to hardware. No real-world fine-tuning. Deployment: months of on-site adaptation → days. Tune into the NVIDIA livestream on 12 August to hear how these companies are closing the sim2real gap: NVIDIA Robotics ~~ ♻️ Join the weekly robotics newsletter, and never miss any news →

Lukas Ziegler

119,091 просмотров • 19 дней назад