Loading video...

Video Failed to Load

Go Home

These are not CGI. Reinforcement learning is so back. When operating on strings, it gives us o3. When operating on physical motors, it gives us a perfect humanoid backflip and a robot creature that out-maneuvers almost every animal on earth. RL is one of the only learning algorithms that...

356,912 views • 1 year ago •via X (Twitter)

11 Comments

Aleksa Gordić 🍿🤖's profile picture
Aleksa Gordić 🍿🤖1 year ago

boston dynamics demo is controlled and we were capable of doing that for years now (e.g. their Atlas robot) but the thing that unitree robot is doing in the natural environment i find it hard to believe this is not cgi? this is a ChatGPT moment for robotics? why isn't this all over the timelines?

Digital Currency's profile picture
Digital Currency2 years ago

From 3D modeling to VR/AR development, our MSc in Metaverse program equips you with the technical skills to excel in the rapidly evolving digital world. Don't miss out—enroll today! #UNIC #MScMetaverse

AI Notkilleveryoneism Memes ⏸️'s profile picture
AI Notkilleveryoneism Memes ⏸️1 year ago

the age of the human generated reward function is over. the age of the AI generated reward function has come.

Abe Murray's profile picture
Abe Murray1 year ago

For folks close to this - how brittle is this? The issue with Boston Dynamics was always the blooper reel and the brittle nature of the reality behind the videos. Is reliability here 95%? 99%? 5x9's? Any insiders have a take on this?

Yuchen Jin's profile picture
Yuchen Jin1 year ago

Unitree is crazy

Chris Paxton's profile picture
Chris Paxton1 year ago

Do you have a source on that BD video being RL?

Shiraz Akmal's profile picture
Shiraz Akmal1 year ago

Reward function

Thanos Variant's profile picture
Thanos Variant1 year ago

That Unitree is so close to being an Autobot

Tianbao Xie's profile picture
Tianbao Xie1 year ago

"Give me a reward function, and I shall move the world." Time to scale on it!!!

H's profile picture
H1 year ago

How often do you have to replace those wheels?

Randall Briggs's profile picture
Randall Briggs1 year ago

What are the reward functions in these cases, would you say?

Related Videos

AI has had exactly two scaling axes that worked so far, and the second one is starting to look finite too the first one was pretraining: with scaling parameters and data, we got world knowledge (i.e. ChatGPT had read enough to know things), but it started saturating a while ago the second one was RL, and people had been doing RL the whole time before that: RLHF is RL but it never scaled far because it was trying to control the exact output, which tokens come out, how the text reads, but you can only push that so far before you’re just polishing RLVR dropped that constraint: giving the model a task, then checking whether the final answer is right, and ignoring everything in between -- so the model does whatever it wants in the middle and only the endpoint gets graded, and that’s much closer to actual RL and it’s what bought us planning and reasoning (arguably, tool use sits around 2.5 on this list -- while useful, it's not a different kind of thing) so one axis gave knowledge, the other gave reasoning, and both of them are one model working alone the next axis is how many models you can get working on the same problem, which is a different kind of axis than the previous two we know that multi-agent RL has always been the harder problem: I spent years in that literature and the gap between single-agent and multi-agent is definitely not incremental -- it’s a whole different class of difficulty! which is also why the derivatives are steep at the start, nobody has picked the easy wins yet... and the thing that gates this multi-agent coordination is communication: models can only coordinate as well as they can exchange information, and right now they do that by writing sentences to each other imagine what could we possibly achieve if we properly open that third axis development by letting models to exchange information in their native "language" without loosing any computational data that they produce during inference

Sasha Malysheva

14,445 views • 1 month ago

This work makes a humanoid robot do simple parkour moves by looking with a depth camera and choosing the right move on the fly. The big deal is that it turns lots of small human moves into long, real-time robot behavior, without hand-coding every transition or retraining for each new course. A humanoid robot is usually good at steady walking, but it often fails when it has to do fast moves like jumping up, vaulting, or rolling, and then keep going to the next obstacle. The hard part is that you cannot easily collect training data for every possible obstacle shape, distance, and mistake, so robots end up learning a few moves that only work in a narrow setup. This work starts from short clips of real human parkour moves, like stepping over, vaulting, climbing, and rolling. It uses motion matching, which is basically a smart “pick the next clip that fits best right now” search, to stitch those short clips into a long, smooth plan that looks like a human doing a whole course. Then it trains a controller with reinforcement learning (RL), which means the robot learns by trial and error to copy that plan while staying balanced and not falling. After training separate expert controllers for different moves, it compresses them into 1 controller that uses only onboard depth sensing and a simple “go this fast in this direction” command. In real tests on a Unitree G1 humanoid, it can clear multiple obstacles in a row, adapt when obstacles get moved, and climb a wall up to 1.25m.

Rohan Paul

37,121 views • 6 months ago

Elon Musk gave the entire entertainment industry its expiration date, and he is the one building the thing that kills it. Musk: “My guess is that we see the first compelling half hour, pure AI show next year.” Next year. A complete show generated entirely by AI. No writers. No actors. No cameras. No sets. No crew. No studio. Just a prompt and enough compute to render a reality that never physically existed. And shows are the easy part. Musk: “I say probably we’re maybe three years away from AI does the whole video game.” A show plays the same way every time. A game has to generate a living world that reacts to every decision in real time across every single frame. That is a fundamentally harder class of problem. And Musk put three years on it. Right now a single AAA title takes seven years and half a billion dollars across thousands of engineers and artists just to ship it. Musk is describing a world where one person types a paragraph and gets something comparable. The entire value proposition of a multi-billion dollar industry lives inside that gap. And it closes in thirty-six months. But the prediction is not the story. The person making it is. This is not an analyst speculating from the sidelines. This is the man building the largest AI compute clusters on the planet. The man who built xAI from zero in under two years. The man stacking hundreds of thousands of GPUs into facilities designed to do exactly what he is describing. When Musk says three years, he is not guessing about what someone else might eventually ship. He is reading you a delivery date off his own roadmap. Every media company on Earth is valued on a single assumption. That quality content is expensive and difficult to produce at scale. That one assumption is the structural foundation underneath every studio, every network, and every publisher in existence. Musk is dismantling it with raw compute. The studios still parading thousand-person production teams are not demonstrating strength. They are advertising the exact cost structure that one person with a prompt and a GPU allocation is about to make irrelevant. And it does not stop at entertainment. If AI can generate an interactive world that responds to human input in real time, it can generate anything. Advertising. Architecture. Training simulations. Product design. Every industry built on humans manually constructing visual experiences frame by frame is sitting on the same countdown Musk just read out loud. Now zoom out. Because this is not just an industry story. For the entire history of human civilization, the distance between imagining a world and actually creating one required thousands of people, millions of hours, and billions of dollars. That distance built Hollywood. That distance built the gaming industry. That distance made content scarce and studios powerful. Musk is collapsing that distance to zero. When the gap between imagining something and it existing disappears, every business model built on the difficulty of creation disappears with it. That is not disruption. That is a full inversion of how human beings create. Musk did not make a casual prediction on that podcast. He told you what he is building. He told you the timeline. And he told you which industries do not survive it. The entertainment industry is still debating whether this future is real. Musk is not part of that debate. He is building. And he just told you the delivery date.

Dustin

22,458 views • 2 months ago

When we say that the aggressor must not receive any reward for the war, so that peace can truly last, everyone must understand – these are not just words. Unfortunately, the world already had the chance to verify this twelve years ago. Russia’s war against Ukraine began with the occupation of our Crimea, and the world effectively turned a blind eye to this. The leaders back then showed little interest in the protests and resistance in Crimea and in Ukraine’s overall sentiment. The world advised Ukraine to stay silent. That is exactly why Putin came to believe he could wage a much larger war and confront the West even more aggressively. Now, every year on February 26 – the Day of Resistance to the Occupation of Crimea – we remember this global lesson, honouring those who did not stay silent and did not falter in the face of Russian aggression, and we insist that holding the aggressor accountable for war is one of the strongest security guarantees – one of the strongest prerequisites for lasting peace. I thank everyone around the world who supports us in this, participates in the work of our Crimea Platform and other international formats that remind the world about Crimea and the meaning of its occupation by Russia. I thank everyone who helps Ukraine resist Russian repression in Crimea, assists us in returning people from captivity, and prevents the occupying regime from consolidating its power. Russian presence on our peninsula serves only war, and nothing else. There must be peace, and therefore Crimea is Ukraine, and the world must always recognize this fact. Qırım evine qaytmalıdır! Glory to Ukraine!

Volodymyr Zelenskyy / Володимир Зеленський

154,120 views • 6 months ago