Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

This is a neural network flying a drone at extremely high speed, beating human champions in FPV drone racing. - Reinforcement learning as a tool is so marvelously versatile. It's able to solve both fast, reactive tasks and slow, deliberate tasks (ChatGPT RLHF). - Trained in large-scale simulation, finetuned...

598,930 görüntüleme • 3 yıl önce •via X (Twitter)

10 Yorum

Jim Fan profil fotoğrafı
Jim Fan3 yıl önce

Original post from the author:

Jeff Holmes profil fotoğrafı
Jeff Holmes3 yıl önce

Who's building the tiny fly and mosquito catching drone for us?

Amy Robinson Sterling profil fotoğrafı
Amy Robinson Sterling3 yıl önce

one step closer to house robots

Per-Anders Edwards profil fotoğrafı
Per-Anders Edwards3 yıl önce

There’s a positive use for deliveries, but there are oh so many more negative uses and you can be sure the DoD and domestic enforcement are going to use every last one of them. So well done that team I guess, you achieved your goal and now you get to reap your reward.

Dave Lavallee profil fotoğrafı
Dave Lavallee3 yıl önce

The implications of this achievement are just scratching the surface of what will be possible in the future. Kudos to the team.

Harri Hakulinen 🇫🇮🇺🇦🇪🇺 profil fotoğrafı
Harri Hakulinen 🇫🇮🇺🇦🇪🇺3 yıl önce

The (near) future of warplanes / wingmans..

𝙀𝙧𝙣𝙖 𝙆𝙧𝙪𝙢𝙗𝙖𝙘𝙝 profil fotoğrafı
𝙀𝙧𝙣𝙖 𝙆𝙧𝙪𝙢𝙗𝙖𝙘𝙝3 yıl önce

When I saw this simulation it reminded me of Harry Potter movie when Harry playing soccer with the broomstick 🧹. Probably NIMBUS1000 is good name for super high speed drone . Just like Harry’s broomsticks.

Lucas Cooper-Bey profil fotoğrafı
Lucas Cooper-Bey3 yıl önce

Well it’s been fun. Hope AGI is stoked on humans

Syrsly (bsky:syrsly.com) profil fotoğrafı
Syrsly (bsky:syrsly.com)3 yıl önce

They're essentially handicapping the race course with standard track shapes instead of random objects. I'm sure it would require a lot more training to handle boxes of different colors, different lighting conditions, rain, triangles... you get the idea.

Φιμπονάτσι profil fotoğrafı
Φιμπονάτσι3 yıl önce

ww3 will be fought with drones/lasers/ai no question

Benzer Videolar

We are thrilled to share our breakthrough research on "Agile Flight from Pixels without State Estimation," to be presented and live-demonstrated at #RSS2024 next week! You heard well: no state estimation means no explicit visual localization, no SLAM, no VIO, and no IMU! Paper: Video (Narrated): Last year, we demonstrated that #ReinforcementLearning (RL) policies could outperform world-champion drone-racing pilots using the same quadrotor hardware; however, unlike human pilots, these policies continuously estimated an explicit state from known gate positions, the camera feed, and inertial measurements (IMU). In this new work, we tackle the challenge of learning vision-based drone racing using an end-to-end reinforcement learning approach that eliminates the need for IMU data or explicit state estimation. Like professional pilots, we go directly from images to control commands. The training is facilitated by an asymmetric actor-critic with access to privileged information. To overcome the computational complexity during image-based RL training, we use an appropriate sensor representation, which can be efficiently simulated during training without rendering images. We achieve agile flight at speeds up to 40 km/h with accelerations up to 2 g's. Although our demonstration focuses on drone racing, we believe that our method has an impact beyond drone racing and can serve as a foundation for future research into real-world applications in structured environments. Besides the paper presentation, we will also give a live demo next Tuesday and Wednesday between and hrs at TU Delft: Reference: Ismail Geles*, Leonard Bauersfeld*, Angel Romero, Jiaxu Xing, Davide Scaramuzza "Demonstrating Agile Flight from Pixels without State Estimation" Robotics: Science and Systems (RSS), 2024. Kudos to Ismail Geles Leonard Bauersfeld Ángel Romero Jiaxu Xing! University of Zurich UZH Science UZH Space Hub Aerial Core AUTOASSESS European Research Council (ERC)

Davide Scaramuzza

28,002 görüntüleme • 2 yıl önce

Can an inexpensive, off-the-shelf IMU be the only sensor to estimate the full state (position, velocity, orientation) of a quadrotor flying through a track at high speed and even be on-pair with vision-based localization? The answer is yes, within certain limitations! In this #RAL2023 paper, we propose a learning-based odometry algorithm that couples a model-based filter driven by the inertial measurements with a learning-based module with access to the control commands. Our system outperforms by a large margin the state-of-the-art visual-inertial odometry (#VIO) algorithms and the state-of-the-art learned-inertial odometry algorithm, #TLIO, for the task of drone racing. Additionally, we show that our system is as accurate as a VIO algorithm that uses a camera to localize to a known map of the racing track. The main limitation of our approach is that it cannot generalize to trajectories that have not been seen at training time. However, in drone racing competitions, the track is known beforehand. Human pilots spend hours or even days of practice on the race track before the competition. Similarly, our system can be trained with the data collected during practice time and deployed during the competition. Future work will investigate how to generalize to trajectories not seen at training time. The code is released! Paper: Video: Code: Kudos to Giovanni Cioffi Leonard Bauersfeld Elia Kaufmann European Research Council (ERC) University of Zurich UZH Science UZH Space Hub NCCR Robotics Aerial Core #RAL2023 #IROS2023 #SLAM

Davide Scaramuzza

37,061 görüntüleme • 3 yıl önce

I don’t know if we live in a Matrix, but I know for sure that robots will spend most of their lives in simulation. Let machines train machines. I’m excited to introduce DexMimicGen, a massive-scale synthetic data generator that enables a humanoid robot to learn complex skills from only a handful of human demonstrations. Yes, as few as 5! DexMimicGen addresses the biggest pain point in robotics: where do we get data? Unlike with LLMs, where vast amounts of texts are readily available, you cannot simply download motor control signals from the internet. So researchers teleoperate the robots to collect motion data via XR headsets. They have to repeat the same skill over and over and over again, because neural nets are data hungry. This is a very slow and uncomfortable process. At NVIDIA, we believe the majority of high-quality tokens for robot foundation models will come from simulation. What DexMimicGen does is to trade GPU compute time for human time. It takes one motion trajectory from human, and multiplies into 1000s of new trajectories. A robot brain trained on this augmented dataset will generalize far better in the real world. Think of DexMimicGen as a learning signal amplifier. It maps a small dataset to a large (de facto infinite) dataset, using physics simulation in the loop. In this way, we free humans from babysitting the bots all day. The future of robot data is generative. The future of the entire robot learning pipeline will also be generative. 🧵

Jim Fan

165,246 görüntüleme • 1 yıl önce

Former Meta Chief AI Scientist Yann LeCun on the three paradigms of machine learning — and why the third is what made ChatGPT possible: Here's each one, and where it breaks. First, supervised learning. You tell the machine the answer. "You show it a picture, let's say of a table, and you tell it this is a table. So it's supervised because you tell it what the correct answer is." Get it wrong, and the machine rewrites itself: "The system computes its output, and if it says something else than table, then it's going to adjust its parameters, its internal structure, so that the output it produces gets closer to the output you want." Repeat at scale and something more than memorisation appears: "Eventually the system will find a way to recognize every image you trained it on, but also images it's never seen that are similar to the one you train it on. This is called a generalization ability." The limit: a human has to supply every single answer. That doesn't scale to the size of the internet. Second, reinforcement learning. You don't give the answer, only a verdict. "You don't tell the system what the correct answer is. You only tell it whether the answer it produced was good or bad." Learning to ride a bike, essentially: "You try to ride a bike and you don't know how to ride the bike and after a while you fall. So you know you did something bad and so you change your strategy a little bit. And eventually you learn how to ride a bike." For years the field assumed this was the closest thing to how animals actually learn. Yann LeCun's verdict: "Now it turns out reinforcement learning is extremely inefficient." It dominates wherever failure is free: "It works really well if you want to train a system to play chess or play go or poker, because you can have the system play millions and millions of games against itself and basically fine-tune itself. But it doesn't really work in the real world." The limit, in one image: "If you want to train a car to drive itself, you're not going to do it with reinforcement learning. It's going to crash thousands of times." On robotics he's careful rather than dismissive: "Reinforcement learning can be part of the solution, but it's not the complete answer. It's not sufficient." Third, self-supervised learning. You tell the machine nothing at all. "And this is what has enabled the recent progress in natural language understanding and chatbots." The strange part is that you stop asking for a task: "You don't train the system to accomplish any particular task. You just train it to basically capture the structure..." The method is deliberate sabotage: "You take a piece of text, you corrupt it in some way, by for example removing some words, and then you train a big neural net to predict the words that are missing." And one narrow version of that trick runs every chatbot on Earth: "A special case of this is that you take a piece of text and the last word in that text is not visible, and so you train the system to predict the last word in that text — and this is the way large language models are trained on." So why did the third one win? Supervised learning needs a human. Reinforcement learning needs a crash. Self-supervised learning needs neither — because the missing word and the correct answer are the same thing. The data grades itself.

Big Brain AI

49,293 görüntüleme • 1 ay önce

New Course: Reinforcement Fine-Tuning LLMs with GRPO! Learn to use reinforcement learning to improve your LLM performance in this short course, built in collaboration with Predibase by Rubrik, and taught by Travis Addair, its Co-Founder and CTO, and Arnav Garg, its Senior Engineer and Machine Learning Lead. Reasoning models have been one of the most important developments in LLMs. Reinforcement Fine-Tuning (RFT) uses rewards to encourage LLMs to find solutions to multi-step reasoning tasks such as solving math problems and debugging code - without needing pre-existing training examples like in traditional supervised fine-tuning. Group Relative Policy Optimization (GRPO) is a reinforcement fine-tuning algorithm gaining rapid adoption. Developed by the DeepSeek team and used to train the R1 reasoning model, GRPO uses reward functions that you can write in Python to assign rewards to model responses. It’s beneficial for tasks with verifiable outcomes and can work well even with fewer than 100 training examples. It can also significantly improve the reasoning ability of smaller LLMs, making applications faster and more cost effective. In this course, you’ll take a technical deep dive into RFT with GRPO. You’ll learn to build reward functions that you can use in the GRPO training process to guide an LLM toward better performance on multi-step reasoning tasks. In detail, you’ll: - Learn when reinforcement fine-tuning is a better fit than supervised fine-tuning, especially for tasks involving multi-step reasoning or limited labeled data. - Understand how GRPO uses programmable reward functions as a more scalable alternative to the human feedback required for other reinforcement learning algorithms, such as RLHF and DPO. - Frame the Wordle game as a reinforcement fine-tuning problem and see how an LLM can learn to plan, analyze feedback, and improve its strategy over time. - Design reward functions that power the reinforcement fine-tuning process. - Learn techniques for evaluating more subjective tasks, such as rating the quality of a text summary, using an LLM as a judge. - Understand why reward hacking happens and how to avoid it by adding penalty functions to discourage undesirable behaviors. - Learn the four key components of the loss calculation in the GRPO algorithm: token probability distribution ratios, advantages, clipping, and KL-divergence. - Launch reinforcement fine-tuning jobs using Predibase’s hosted training services. By the end of this course, you’ll be able to build and fine-tune LLMs using reinforcement learning to improve reasoning without relying on large labeled datasets or subjective human feedback. Please sign up here:

Andrew Ng

86,697 görüntüleme • 1 yıl önce

Can GPT-4 teach a robot hand to do pen spinning tricks better than you do? I'm excited to announce Eureka, an open-ended agent that designs reward functions for robot dexterity at super-human level. It’s like Voyager in the space of a physics simulator API! Eureka bridges the gap between high-level reasoning (coding) and low-level motor control. It is a “hybrid-gradient architecture”: a black box, inference-only LLM instructs a white box, learnable neural network. The outer loop runs GPT-4 to refine the reward function (gradient-free), while the inner loop runs reinforcement learning to train a robot controller (gradient-based). We are able to scale up Eureka thanks to IsaacGym, a GPU-accelerated physics simulator that speeds up reality by 1000x. On a benchmark suite of 29 tasks across 10 robots, Eureka rewards outperform expert human-written ones on 83% of the tasks by 52% improvement margin on average. We are surprised that Eureka is able to learn pen spinning tricks, which are very difficult even for CGI artists to animate frame by frame! Eureka also enables a new form of in-context RLHF, which is able to incorporate a human operator’s feedback in natural language to steer and align the reward functions. It can serve as a powerful co-pilot for robot engineers to design sophisticated motor behaviors. As usual, we open-source everything! Welcome you all to check out our video gallery and try the codebase today: Paper: Code: Deep dive with me: 🧵

Jim Fan

2,677,701 görüntüleme • 2 yıl önce

Let's reverse engineer Disney's adorable, lifelike robot! I couldn't find a whitepaper, but this is how I think it's trained: 1. The emotional behaviors are curated by Disney animation artists, keyframe by keyframe. But it cannot be "rendered" directly on the robot because it doesn't take into account the complex real-world physics. 2. Reinforcement learning (RL) is a great tool for training low-level robot controllers. RL needs a reward function to optimize, and it's typically a task reward (e.g. walk in a straight line as fast as possible). The problem is that RL doesn't know what counts as "natural behavior", and often produces weird-looking body postures that somehow still maximize the reward. This is a human alignment problem just like ChatGPT. 3. Enters Adversarial Motion Prior (AMP): a technique that learns the human preference by training a classifier on what we consider "emotional & cute". In GAN literature, this is called a discriminator. Disney artists are good at creating such a dataset. You can then add AMP as an auxiliary reward in simulation to nudge the robot towards desired behaviors. AMP was developed by Peng et al. 2021 and Escontrela et al. 2022. 4. Add lots of data augmentation to make the controller robust to physical disturbances. In RL, it's called "domain randomization". This is a very powerful technique that bridges the gap between simulator and reality. Previously, OpenAI used domain randomization to train a 5-finger robot hand to manipulate a Rubik's Cube: IEEE news article gave hints about the pipeline: Finally, praying for world peace 🙏. I hope robotics like this will bring more joy to the world.

Jim Fan

314,807 görüntüleme • 3 yıl önce

Today, we're joined by Nikita Rudin, co-founder and CEO of Flexion to discuss the gap between current robotic capabilities and what’s required to deploy fully autonomous robots in the real world. Nikita explains how reinforcement learning and simulation have driven rapid progress in robot locomotion—and why locomotion is still far from “solved.” We dig into the sim2real gap, and how adding visual inputs introduces noise and significantly complicates sim-to-real transfer. We also explore the debate between end-to-end models and modular approaches, and why separating locomotion, planning, and semantics remains a pragmatic approach today. Nikita also introduces the concept of "real-to-sim", which uses real-world data to refine simulation parameters for higher fidelity training, discusses how reinforcement learning, imitation learning, and teleoperation data are combined to train robust policies for both quadruped and humanoid robots, and introduces Flexion's hierarchical approach that utilizes pre-trained Vision-Language Models (VLMs) for high-level task orchestration with Vision-Language-Action (VLA) models and low-level whole-body trackers. Finally, Nikita shares the behind-the-scenes in humanoid robot demos, his take on reinforcement learning in simulation versus the real world, the nuances of reward tuning, and offers practical advice for researchers and practitioners looking to get started in robotics today. 🗒️ For the full list of resources for this episode, visit the show notes page: 📖 CHAPTERS =============================== 00:00 - Introduction 04:07 - Is robot locomotion solved? 06:04 - Sim-to-real gap 08:58 - Adding semantics to policies 09:42 - Modular vs end-to-end architectures 10:29 - Planner model 12:21 - Adapting RL techniques from quadrupeds to humanoids 15:39 - Behind robot demos 18:09 - Humanoid robots in home environments 22:03 - Training approach 23:56 - VLA models 27:59 - Closing the sim-to-real gap 32:55 - Task orchestration using VLMs 36:38 - Tool use 38:10 - Model hierarchy 43:37 - Simulator versus simulation environment 44:57 - Combining imitation learning and reinforcement learning 46:42 - RL in real world versus RL in simulation 52:58 - Reward tuning and value functions in robotics 56:38 - Predictions 1:00:10 - Humanoids, quadropeds, and wheeled platforms 1:02:45 - Advice, recommended robot kits, and community pla

The TWIML AI Podcast

22,592 görüntüleme • 8 ay önce