Loading video...

Video Failed to Load

Go Home

Physics-based simulation tools like DFT and FEP are being repurposed for dataset development. Once manually run and reviewed by experts, these tools are now producing tens of millions of synthetic data points without close human oversight. These techniques are not without their issues. Studies have shown that in certain...

10,446 views • 10 days ago •via X (Twitter)

2 Comments

ธีรภัทร ไวยครุฑ's profile picture
ธีรภัทร ไวยครุฑ10 days ago

The societal implications of reinforcement learning depend heavily on transparency, accountability, and governance in real-world applications.

Rachel St. Clair's profile picture
Rachel St. Clair10 days ago

Same failure mode shows up with LLM-generated synthetic data for post-training: once nobody's spot-checking, the generator's own biases compound across millions of samples instead of just the target signal you wanted.

Related Videos

Scale alone is not enough for AI data. Quality and complexity are equally critical. Excited to support all of these for LLM developers with Snorkel AI Data-as-a-Service, and to share our new leaderboard! — Our decade-plus of research and work in AI data has a simple point: scale alone is not enough. AI success is all about the quality, complexity, and distribution of data—in addition to volume. We’re excited to be powering leading LLM developers with Snorkel AI Expert Data-as-a-Service, our white glove service for custom, expert-level AI datasets—and to now preview some of what we’re building via our new Expert Data Leaderboard (🔗 in 🧵) + upcoming OSS dataset releases! Snorkel Expert Data-as-a-Service is built to meet the rapidly evolving data needs of the agentic AI world—where success is built on the quality, complexity, and distribution of datasets, in addition to size and scale. This kind of high-quality, frontier AI data can only come from a union of technology and human expertise. With Snorkel Expert Data-as-a-Service, we’re powering frontier LLM developers across agentic, expert knowledge, reasoning, coding, multi-modal, and other task types via the combination of these two key components: - (1) The Snorkel Expert Network: A global team of subject matter experts focused wholly on specialized knowledge–spanning thousands of topics in STEM/academic, vertical/professional, and consumer/lifestyle domains. - (2) Snorkel AI Data Development Platform: Our unique programmatic data curation and quality control platform, accelerating and improving expert authoring and review through principled techniques developed over the last decade of R&D. Now: we’re incredibly excited to showcase some of the power of Snorkel Expert Data-as-a-Service via the new Snorkel Leaderboard—putting frontier models to the test in complex, agentic, and reasoning settings inspired by real industry scenarios (not esoteric puzzles)! We’ll be releasing new leaderboards and accompanying expert-verified open source datasets (coming soon!) regularly. To start, we’re sharing three initial ones in preview: - SnorkelFinance: Q&A over financial documents requiring agentic tool-calling and reasoning - SnorkelUnderwrite: Agentic insurance tasks requiring industry-specific reasoning and tool use - SnorkelSequences: Mathematical tasks requiring compositional multi-step reasoning

Alex Ratner

495,851 views • 1 year ago

I don’t know if we live in a Matrix, but I know for sure that robots will spend most of their lives in simulation. Let machines train machines. I’m excited to introduce DexMimicGen, a massive-scale synthetic data generator that enables a humanoid robot to learn complex skills from only a handful of human demonstrations. Yes, as few as 5! DexMimicGen addresses the biggest pain point in robotics: where do we get data? Unlike with LLMs, where vast amounts of texts are readily available, you cannot simply download motor control signals from the internet. So researchers teleoperate the robots to collect motion data via XR headsets. They have to repeat the same skill over and over and over again, because neural nets are data hungry. This is a very slow and uncomfortable process. At NVIDIA, we believe the majority of high-quality tokens for robot foundation models will come from simulation. What DexMimicGen does is to trade GPU compute time for human time. It takes one motion trajectory from human, and multiplies into 1000s of new trajectories. A robot brain trained on this augmented dataset will generalize far better in the real world. Think of DexMimicGen as a learning signal amplifier. It maps a small dataset to a large (de facto infinite) dataset, using physics simulation in the loop. In this way, we free humans from babysitting the bots all day. The future of robot data is generative. The future of the entire robot learning pipeline will also be generative. 🧵

Jim Fan

165,246 views • 1 year ago

Exciting updates on Project GR00T! We discover a systematic way to scale up robot data, tackling the most painful pain point in robotics. The idea is simple: human collects demonstration on a real robot, and we multiply that data 1000x or more in simulation. Let’s break it down: 1. We use Apple Vision Pro (yes!!) to give the human operator first person control of the humanoid. Vision Pro parses human hand pose and retargets the motion to the robot hand, all in real time. From the human’s point of view, they are immersed in another body like the Avatar. Teleoperation is slow and time-consuming, but we can afford to collect a small amount of data. 2. We use RoboCasa, a generative simulation framework, to multiply the demonstration data by varying the visual appearance and layout of the environment. In Jensen’s keynote video below, the humanoid is now placing the cup in hundreds of kitchens with a huge diversity of textures, furniture, and object placement. We only have 1 physical kitchen at the GEAR Lab in NVIDIA HQ, but we can conjure up infinite ones in simulation. 3. Finally, we apply MimicGen, a technique to multiply the above data even more by varying the *motion* of the robot. MimicGen generates vast number of new action trajectories based on the original human data, and filters out failed ones (e.g. those that drop the cup) to form a much larger dataset. To sum up, given 1 human trajectory with Vision Pro -> RoboCasa produces N (varying visuals) -> MimicGen further augments to NxM (varying motions). This is the way to trade compute for expensive human data by GPU-accelerated simulation. A while ago, I mentioned that teleoperation is fundamentally not scalable, because we are always limited by 24 hrs/robot/day in the world of atoms. Our new GR00T synthetic data pipeline breaks this barrier in the world of bits. Scaling has been so much fun for LLMs, and it's finally our turn to have fun in robotics! We are building tools to enable everyone in the ecosystem to scale up with us. Links in thread:

Jim Fan

364,670 views • 2 years ago

My name is Chris Elston, and I’m here to speak for children being irreversibly harmed by health professionals practicing what is falsely referred to as gender-affirming care. It is not caring to stop the development of children with puberty blockers. These are repurposed chemical castration and cancer drugs. It is not caring to alter their development with opposite-sex hormones. Every systematic review of the scientific literature shows that children are being harmed, and that scientific rigour is non-existent. Activist organizations pushing ideology have hijacked the medical community, and what is being done is nothing more than a live, unregulated experiment on kids. It is one of the great deceptions of our time to teach children that they might have been born in the wrong bodies! Affirming care would be to tell them that they are beautiful just as they are. No drugs or scalpels needed! Instead, children are being sterilized, and are having healthy body parts cut off. Overwhelmingly, these kids are autistic, have other mental health comorbidities, and many have suffered trauma or abuse. Girls as young as 12 are having their breasts removed, and 16-year-old boys are being castrated. State encroachment on parental rights is worsening outcomes. Here in Geneva, a child was taken from her parents because they refuse to transition her. Children have the human right to grow up with their bodies intact. It is time the Member States of the United Nations use the resolutions, procedures, and other tools at their disposal to stop this child abuse!

Billboard Chris 🌎

46,073 views • 6 months ago

On this Transgender Day of Remembrance, we remember my speech at the United Nations Human Rights Council. My name is Chris Elston, and I am to here to speak for children being irreversibly harmed by health professionals practicing what is falsely referred to as gender-affirming care. It is not caring to stop the development of children with puberty blockers. These are repurposed chemical castration and cancer drugs. It is not caring to alter their development with opposite-sex hormones. Every systematic review of the scientific literature shows that children are being harmed, and that scientific rigour is non-existent. Activist organizations pushing ideology have hijacked the medical community, and what is being done is nothing more than a live, unregulated experiment on kids. It is one of the great deceptions of our time to teach children that they might have been born in the wrong bodies! Affirming care would be to tell them that they are beautiful just as they are. No drugs or scalpels needed! Instead, children are being sterilized, and are having healthy body parts cut off. Overwhelmingly, these kids are autistic, have other mental health comorbidities, and many have suffered trauma or abuse. Girls as young as 12 are having their breasts removed, and 16-year-old boys are being castrated. State encroachment on parental rights is worsening outcomes. Here in Geneva, a child was taken from her parents because they refuse to transition her. Children have the human right to grow up with their bodies intact. It is time the Member States of the United Nations use the resolutions, procedures, and other tools at their disposal to stop this child abuse!

Billboard Chris 🌎

46,511 views • 10 months ago

Physics-based Motion Retargeting from Sparse Inputs paper page: Avatars are important to create interactive and immersive experiences in virtual worlds. One challenge in animating these characters to mimic a user's motion is that commercial AR/VR products consist only of a headset and controllers, providing very limited sensor data of the user's pose. Another challenge is that an avatar might have a different skeleton structure than a human and the mapping between them is unclear. In this work we address both of these challenges. We introduce a method to retarget motions in real-time from sparse human sensor data to characters of various morphologies. Our method uses reinforcement learning to train a policy to control characters in a physics simulator. We only require human motion capture data for training, without relying on artist-generated animations for each avatar. This allows us to use large motion capture datasets to train general policies that can track unseen users from real and sparse data in real-time. We demonstrate the feasibility of our approach on three characters with different skeleton structure: a dinosaur, a mouse-like creature and a human. We show that the avatar poses often match the user surprisingly well, despite having no sensor information of the lower body available. We discuss and ablate the important components in our framework, specifically the kinematic retargeting step, the imitation, contact and action reward as well as our asymmetric actor-critic observations. We further explore the robustness of our method in a variety of settings including unbalancing, dancing and sports motions.

AK

106,527 views • 3 years ago

Experiments in progress. The one on the right has been learning for ~3 hours, the one in the middle for ~1 hour, and the one on the left just started a few minutes ago. The initial motivation for making the physical Atari was just to commit ourselves to a subset of algorithms that can make progress in this setup. This commitment rules out algorithms that require billions of samples to learn (or worse, require multiple environments running in parallel). Atari games are simple enough that we should be able to show learning on them in a short amount of time with no prior knowledge. Since then, I've realized that this setup is also a good way to compare different paradigms in robotics in a principled way. These paradigms are sim2real, learning from tele-operated data, and learning directly on the robots. So far, I have observed that getting sim2real to work reliably is hard. It requires tweaks that don't scale. Policies that can play perfectly in simulation fall apart because of latencies and the messiness of the real world. These aspects could be modeled to improve the simulation, but not without sinking significant human engineering hours. I have higher hopes for learning from tele-operated data, but that requires a human to learn the task first. These experiments are on my to-do list. I have to learn to play some of the games well through the robot. I’m half-decent at playing Pong and Ms Pacman now. Learning directly on robots is looking like the most promising approach. This approach takes away pesky distribution shifts and makes it possible to have algorithms that continually improve with more data and time without any human intervention. It feels great to let experiments run overnight and wake up to find improved policies. With learning on robots, I should, in principle, be able to go on a long vacation and come back to find better policies for complex tasks beyond Atari games. Whether that is possible with current learning algorithms is a different question.

Khurram Javed

52,110 views • 10 months ago