Video wird geladen...
Video konnte nicht geladen werden
A big part of scaling robot learning to solve real-world problems is that we somehow need to get enough diverse, high-quality data to train our robots to perform useful things. GPT and its fellow large language models were bootstrapped and proved out on a massive dataset of real-world language... show more
20,486 Aufrufe • vor 1 Jahr •via X (Twitter)
11 Kommentare

Learning Robotic Manipulation from Simulations A comparison of a few recent works on sim-to-real robot manipulation that I liked

From semiconductors to data centers, AIS targets the critical components behind AI's exponential growth. Capture potential returns from this transformative technology sector.

uncontested Isaac superiority 🫡 (very excited for warp too)

what are you thoughts on a model that you just shovel in alot multimodal data and how much of a ratio you’ll need to get good performance? for example recent work from @physical_int they made an architecture where they can predict web data and predict actions.

Sim to real is definitely not the only way to do this I wrote a previous post that mentioned this; I wrote the two together so they reference the same paper on how much sim data you "need": I think it's worth a deeper look. but my guess is that sim and real data are fulfilling different needs: - sim data is generally very good at capturing robot planning and kinematics - video data is really good at capturing semantics and information about the world - robot data captures everything but isn't diverse enough and is too expensive (although people are trying to change that)

Excited for real2sim to get easier! Just using your phone + lidar to capture your workspace to port to simulation could be a game-changer for sim2real. Learned simulators could maybe bridge that gap as well

What do you think about the sim2real transfer strategies & trying to scale up data that way, versus the strategies of training on real, non-robot data, which is readily available? I'm thinking about things like DexMachina learning from human ego-centric demonstration (could put smart glasses on people doing their everyday jobs) or Meta's V-JEPA 2 model that relies on the massive corpus of 'things happening in the world' to build a foundation model that has physical understanding.

@chris_j_paxton I'm building browser-native generalized physics simulators and would love to chat about async agentic acceleration and scaling sim data. The current pipeline is incredibly fragmented.

Yeah interesting!

yes, your "probably" is correct: we modeled the table geometry as a collision in dextrah (the original one, wasn't part of the rgb extension)

the comparison doesn't seem to hold the amount of info u can infer from the universe is literally (uncountably) infinite, and the amount of knowledge we already discover is huge already, hence the necessity for gpt to have access to huge amount of data but for physical task ?..
