Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Unitree burned $4M to make their G1 humanoid beat pro drummers in real-time rhythm The result is terrifyingly good, mastering micro-second timing required training neural motor control from scratch. Instead of pre-programmed loops, the robot processes dynamic feedback on the fly to hit exact percussion velocity and timing. Why...

45,898 görüntüleme • 12 gün önce •via X (Twitter)

0 Yorum

Yorum bulunmuyor

Orijinal gönderinin yorumları burada görünecek

Benzer Videolar

Are you watching the Chinese New Year Gala? The Robot Kungfu show is mind blowing!!! They just executed a coordinated martial arts routine with spatial precision, rhythm control, and dynamic balance adjustments in real time. Kung fu, one of China’s most iconic traditional art forms , performed by machines built with cutting-edge AI control systems, advanced actuators, and high-speed feedback loops. Ancient discipline meets algorithmic precision. Last year, humanoid robots stepped onto the Spring Festival Gala stage for the first time. This year, they held synchronized kung fu stances with balance that would humble half of us after leg day. And they did it live!!! On the most-watched television event on the planet. The progress in just one year is magical. That’s what we call China speed. What makes it even sweeter is where this happened. I love how the progress is integrated in culture. In celebration. In a Lunar New Year gala watched by hundreds of millions. It’s music to my ears. The robots didn’t look like they were “trying” anymore. They looked like they belonged. Their joint articulation was smoother. Their formation timing tighter. Their balance recovery almost elegant. Their choreography is expressive. That’s what happens when AI models improve, control systems get smarter, hardware stabilizes, and iteration cycles compress. One year in robotics today is not the same as one year ten years ago. It’s compounding. If this is what 12 months looks like, imagine 36. The Chinese New Year Robot Kungfu Gala is just futuristic. It was quite the statement! The future is getting better very, very fast. It was so beautiful to watch. What do you think?

Evrim Kanbur

1,557,133 görüntüleme • 6 ay önce

Not a preplanned motion sequence. A robot deciding mid-jump what to do next. [📍 paper + demo] Researchers just showed a humanoid doing real parkour using only onboard perception. No motion script, no fixed obstacle layout. The system is called Perceptive Humanoid Parkour (PHP). Instead of memorizing a path, the robot reads depth from its cameras and continuously chooses actions. Step, vault, climb, or roll depending on what geometry appears in front of it. To make that possible, they combine three ideas: First, they stitch together human motion clips into long movement references so the robot learns fluid transitions instead of isolated tricks. Second, they train tracking policies with reinforcement learning so contacts land at the right time and the robot keeps balance during dynamic moves. Finally, everything is distilled into one perception policy that runs directly from depth input to action selection. The result on a Unitree G1: about 3 m/s vaults wall climbs up to 1.25 m nearly one minute continuous obstacle traversal adapting when obstacles move What matters is not the tricks. It is the shift in capability. Earlier humanoids executed motions. This one navigates situations. Once robots react to geometry instead of replaying trajectories, environments stop needing to be predictable. Warehouses, homes, and outdoors suddenly become the same problem. Thanks for sharing, Zhen Wu! Paper + demo: ——— Weekly robotics and AI insights. Subscribe free:

Ilir Aliu

22,080 görüntüleme • 6 ay önce

Announcing DreamDojo: our open-source, interactive world model that takes robot motor controls and generates the future in pixels. No engine, no meshes, no hand-authored dynamics. It's Simulation 2.0. Time for robotics to take the bitter lesson pill. Real-world robot learning is bottlenecked by time, wear, safety, and resets. If we want Physical AI to move at pretraining speed, we need a simulator that adapts to pretraining scale with as little human engineering as possible. Our key insights: (1) human egocentric videos are a scalable source of first-person physics; (2) latent actions make them "robot-readable" across different hardware; (3) real-time inference unlocks live teleop, policy eval, and test-time planning *inside* a dream. We pre-train on 44K hours of human videos: cheap, abundant, and collected with zero robot-in-the-loop. Humans have already explored the combinatorics: we grasp, pour, fold, assemble, fail, retry—across cluttered scenes, shifting viewpoints, changing light, and hour-long task chains—at a scale no robot fleet could match. The missing piece: these videos have no action labels. So we introduce latent actions: a unified representation inferred directly from videos that captures "what changed between world states" without knowing the underlying hardware. This lets us train on any first-person video as if it came with motor commands attached. As a result, DreamDojo generalizes zero-shot to objects and environments never seen in any robot training set, because humans saw them first. Next, we post-train onto each robot to fit its specific hardware. Think of it as separating "how the world looks and behaves" from "how this particular robot actuates." The base model follows the general physical rules, then "snaps onto" the robot's unique mechanics. It's kind of like loading a new character and scene assets into Unreal Engine, but done through gradient descent and generalizes far beyond the post-training dataset. A world simulator is only useful if it runs fast enough to close the loop. We train a real-time version of DreamDojo that runs at 10 FPS, stable for over a minute of continuous rollout. This unlocks exciting possibilities: - Live teleoperation *inside* a dream. Connect a VR controller, stream actions into DreamDojo, and teleop a virtual robot in real time. We demo this on Unitree G1 with a PICO headset and one RTX 5090. - Policy evaluation. You can benchmark a policy checkpoint in DreamDojo instead of the real world. The simulated success rates strongly correlate with real-world results - accurate enough to rank checkpoints without burning a single motor. - Model-based planning. Sample multiple action proposals → simulate them all in parallel → pick the best future. Gains +17% real-world success out of the box on a fruit packing task. We open-source everything!! Weights, code, post-training dataset, eval set, and whitepaper with tons of details to reproduce. DreamDojo is based on NVIDIA Cosmos, which is open-weight too. 2026 is the year of World Models for physical AI. We want you to build with us. Happy scaling! Links in thread:

Jim Fan

227,634 görüntüleme • 6 ay önce

Bio inspired Hebbian probabilistic network learns in less than 5 minutes from a super sparse single reward per episode! also has imitation learning (manual control) system has 3 parallel competing networks which get sensory input from a 360 vision (27-direction sensory neuron array) link to code in comment each sub-network is responsible for a single motor action: forward, left and right. at each step whichever section has most neurons firing wins neurons fire probabilistically and mark themselves with a time-decay tag which happens when a neuron fires and diminishes with time. you can see this " tag countdown" on each neuron when a reward is attained(eating the cheese) eligible connections gets strengthened I included 2 runs in the video first was 15 minutes in real time and second was 5 minutes. red plot is the rolling average of last 10 time to cheese. it is really not possible for agent to achieve full control due to probabilistic neural firing. that is why it has to learn while jittering all over the place, which in itself is interesting in manual mode you can guide the cheese by stimulating its motor control networks ( still probabilistically ) and the rewards will still work ✅ Biologically Plausible Features: Stochastic firing (neurons in the brain fire probabilistically) Reward-based learning (dopamine-like neuromodulation) Hebbian plasticity (well-established biological mechanism) Eligibility traces (biological neurons have temporal credit assignment) Sparse sensory encoding (similar to place cells, grid cells) Competitive action selection (basal ganglia architecture) No backpropagation (which is biologically implausible) ❌ Missing Biological Features: No recurrent connections (real brains have extensive feedback loops) No inhibitory neurons (GABAergic neurons are ~20% of cortex) No spike timing (simplified from true spiking dynamics) Uniform layer structure (biological networks are more heterogeneous) Simple weight updates (real synaptic plasticity is more complex)

echo.hive

33,638 görüntüleme • 10 ay önce

Everything Elon said about Optimus at the All-In Summit today: • We’re finalizing the design of Optimus v3. That release is going to be a very remarkable robot. It will have manual dexterity comparable to a human, meaning a very complex hand, an AI mind that can navigate and comprehend reality, and will be made in very high volume. • Other robotics companies are missing those three very hard things. • I spend more mental cycles on Optimus than any other single thing. Solving real-world AI, all of the electrical-mechanical issues, the supply chain, and production challenges. • There is no supply chain for humanoid robots, so it has to be created from scratch, which requires a lot of vertical integration. None of the actuators in Optimus are available from an existing supply chain. • I think if successful, Optimus would be the biggest product ever. • The marginal cost of production, once we hit a million units per year, will probably be around $20,000. It depends on how much we spend on the AI chip in the robot, and we’ll need to achieve a lot of efficiencies in the actuators—26 actuators per arm (26 motors, gearboxes, and power electronics). The AI chip might cost $5,000 or $6,000, maybe more. At 1 million units a year, production cost will be $20,000, maybe $25,000. Price will be a function of demand. • Human hands have evolved to be incredibly sophisticated machines. Hands are a very first instrument. You can swing a baseball bat, thread a needle, play a piano or violin, and assemble a car. Hands are incredibly versatile instruments. Most of the muscles of the hands are actually in the forearm, and the hand is almost like a puppet. Human tendon evolution is incredibly good. The human hand has 27 or 28 degrees of freedom, depending on how you count it; it’s amazing. • In order to create a robot that can be a generalized humanoid, you must solve the “hands problem.” • Even though there are 10,000 to 20,000 electric motors out there, we couldn’t buy the actuators for any amount of money. We had to design every electric motor, gearbox, and controlling electronics from scratch, from first principles of physics. • Optimus is harder than developing any previous Tesla product, but not harder than Starship. • Right now, we’re struggling with the final design of the hardware, primarily the hand. The hands and forearm are the majority of the engineering difficulty of the entire robot. • If you want to do all the things that a human can do, it turns out you need a humanoid robot. If you want to do a subset, that’s much easier. Humans evolved to the shape and capability that we have for a good reason. There is value to having four fingers and a thumb; even the pinky is quite useful. Toes are much more of a question mark. • The AI5 inference chip will be 40 times better than AI4 by some measures. We know the limiting factors of the chip because the AI software and hardware teams work so closely. Effectively, the Tesla AI hardware and software teams are co-designing the chip. • The Softmax function on AI4 takes 40 steps in emulation mode, which will take only a few steps in AI5 natively. AI5 will easily handle mixed precision. • In terms of nominal raw compute, the AI5 inference chip has 8 times more compute, 9 times more memory, and 5 times more memory bandwidth compared to AI4. Because we’re addressing some core limitations and optimizations at the silicon level, we’re able to realize 40x improvements.

The Humanoid Hub

239,049 görüntüleme • 11 ay önce

Imagine controlling a real robot from your home… no money, no experience needed. Sounds crazy, right? But it’s already possible. BitRobot 🦾 is building the world’s first open robotics lab powered by crypto incentives. Instead of one company doing everything, it connects people from all over the world to work together on real robotics and AI tasks. The network is made up of specialized subnets, each focused on different missions from collecting real-world data with robots to developing humanoid robots for everyday use. What makes it powerful? It uses crypto rewards to coordinate global resources like compute power, robot fleets, teleoperation time, and even human effort. This allows BitRobot to scale much faster than traditional labs. Now here’s the best part 👇 The easiest way to get involved right now is through TeleArms. You don’t need: – a robot – engineering skills – or any investment – Hardware All you need is a laptop and an internet connection. From your home, you can remotely control a real robotic arm inside BitRobot’s lab using your keyboard or mouse to pick up, move, and place objects. Every action you take helps generate real-world data that trains the next generation of AI to perform useful physical tasks. So you’re not just playing with a robot… You’re actually helping build the future of AI. I’ve been talking about BitRobot for a while, and now TeleArms is live! You can control a real robotic arm from home, but it’s in a private beta with limited access. I’m now an ambassador for BitRobot Network. I’m giving 4 exclusive access codes to my community so they can experience it too. A lot of people want to experience this, but since it’s limited, I decided to do a random giveaway. To participate in this giveaway : 1. Join the BitRobot Network Discord (Link in comments) 2. Come back to this post and comment below, explaining why you want to join TeleArms and how you plan to contribute. Note : Winner will be announced in the last 7 days. Once you do that, you’ll be in the running for one of the codes! Good luck, and I can’t wait to see your ideas!

Apurba.Eth

36,289 görüntüleme • 5 ay önce

Ex Machina is no longer sci-fi. China has finally built it. The company is AheadForm, founded in Shanghai. The product is the world's most hyper-realistic robotic face. Silicone skin you can't tell from human, 25 micro motors hidden underneath pulling the face into real expressions. And RGB cameras embedded inside the pupils so when it looks at you, it actually sees you from where its eyes are. They raised $28.5M to "give AI a head," which is also where the name comes from. AheadForm = a head form. This is the opposite of where everyone else in robotics is focused. Unitree, Figure, Tesla, Boston Dynamics: all about the body. AheadForm chose the face because they think trust is the harder problem to solve, and trust gets decided at the face. The reason nobody else has tried this is the "uncanny valley." It's the creepy zone where a robot looks almost human but not quite, and looking at it just feels wrong even when you can't say why. Most roboticists believed no amount of engineering could make a face realistic enough to escape it. So they gave up and kept robots cartoonish on purpose: big anime eyes, exaggerated features, clearly synthetic. But AheadForm decided to treat it as an engineering bug instead. Add enough motors, tune the silicone, fix the timing, the valley closes. And they're pulling it off. A few crazy details about how this actually works: 1. The robot learns its own face in a mirror. You put it in front of a camera, let it fire every motor randomly, and it watches what its face does and builds an internal map of "if I send command X to motor Y, my eyebrow does this." Same exact process a human baby uses staring into a mirror. The robot teaches itself who it is by experimenting. 2. It predicts your smile 839 milliseconds before you smile. By watching the micro-tells in your face that precede a smile, the robot starts smiling 0.8 seconds ahead, so its smile lands at the same moment yours does. Most robot mimicry happens half a second late, which is exactly why it always feels artificial. 3. The pupils are the cameras. When the robot makes eye contact, the gaze and the sensor are the same physical thing. Most humanoid robots stick the camera on the forehead or chest, so they aren't actually looking at you when their eyes are pointed at you. 4. The founder, Yuhang Hu, did his PhD at Columbia under Hod Lipson. Lipson is the guy who in 2006 built a four-legged robot that figured out it had four legs by experimenting with its own movement, nobody told it the body shape, it discovered it. He has spent 25 years trying to build machines that know what they are. AheadForm is that 25-year research arc productized. 5. NetEase Games already paid them to physically embody a fantasy video game character. That opens up a brand-new category: robotics as the physical embodiment of fictional IP. Every character-rich studio, Disney, Riot, Hoyoverse, Pokemon, Netflix, now has a question to answer about when their characters get bodies. AheadForm believes whoever ships the first robot you'd actually want around your family wins. That's the bet behind the most realistic robot face on earth.

Ole Lehmann

537,226 görüntüleme • 3 ay önce

Chinese robotics company Astribot released their latest World-Action Model (WAM), Lumo-2. Technical breakdown: - based on a frozen 🥶 Qwen-3.5 4B VLM - trained in 3 progressive stages: 1. Action is aligned with latent world dynamics (an abstract representation of action). Real-world actions are anchored to physical constraints, while the latent space is guided to focus on motion-relevant changes. This bidirectional relationship makes the model physically grounded -> critical for a world model. 2. Action is aligned with vision and language. Reusing the vision backbone and action encoder from the frozen VLM, the authors add a custom vocabulary (for new actions), a semantic module, an action decoder, and an action projector. This aligns the (new) action representations with the (existing) vision-language semantic space. Most importantly: it builds a direct mapping from natural-language instructions to motor execution. 3. End-to-end training on language, video, and robot data. Only the new modules (everything outside the frozen backbone) are trained end-to-end across temporal reasoning, physical understanding, long-horizon, and dexterous manipulation. At the end of the day, Lumo-2 is not the best on benchmarks, but that's not the point. What's genuinely new: - a way to combine latent world modeling and action generation through progressive alignment - a physically-grounded latent dynamics space - it lifts performance on unseen objects using un-annotated human egocentric video + Vision Pro captures, no special transfer algorithm needed Why it matters: - the whole model is thin trainable adapters (semantic module, action decoder/projector) on a frozen 4B backbone (cheap) - that scale is suited for real-time embedded inference (~2.71× decode speedup, no accuracy loss) - its real moat is long-horizon execution, where the added temporal memory pays off far more than on any other task As a result, this robot can now make your latte (5x sped up video):

Léo

32,331 görüntüleme • 1 ay önce

i don't think people realize what's happening in Chinese robotics. this one manufacturer might be the most impressive AND most concerning company on Earth right now let me explain... Unitree Robotics sells a humanoid robot for $5,900. their robot dog costs $1,600 (Boston Dynamics charges $74,500 for theirs for context). you can literally buy these on Amazon today. so obviously the first question is: how is that even possible? the answer starts with a guy who couldn't pass his English exam. Wang Xingxing grew up in Zhejiang province. for his master's thesis, he decided to build a quadruped robot. budget: about $3,000. for context, $3,000 for this kinda robot is nothing. off-the-shelf servo motors alone would've eaten that twice over. so Wang did the only thing he could: he designed and machined every single component himself. motors, joints, controllers, the frame. all of it. the resulting robot was janky and imperfect. but it worked. and the video went viral globally. after graduating he joined DJI. but he quit after two months, and this is 2016, when DJI was arguably the hottest hardware company in China. walking away from that with no money to start a robotics company is a... specific kind of stubborn. he launches Unitree with $280K from a single angel investor. tiny office in Hangzhou. 50 square meters. but the money runs out fast. he can't make payroll for three years. the company almost dies in 2017. but emergency government funding arrives with days to spare. he survives, barely, and keeps building. this is where it gets really fascinating IMO. this founding constraint, building everything yourself because you literally cannot afford to buy parts, never went away. even after funding rounds started landing. even after revenue kicked in. it just became the company's permanent DNA. Unitree now manufactures 90%+ of its core components in-house. motors, reducers, controllers, encoders, LiDAR, etc the founder's $3,000 robot thesis ended up being an architectural decision that turned out to be structurally superior. think about what that means in practice. Boston Dynamics needs a better motor? they negotiate with a supplier, wait on lead times, qualify the part. but when Unitree needs one, they design theirs internally and have a new version in production within weeks. that gap compounds every cycle. Unitree shipped three separate humanoid platforms in 18 months. Figure AI has shipped one. Tesla has shipped zero commercially. the results are getting hard to dismiss. 23,700 robot dogs shipped in 2024 (roughly 70% of the entire global market). 7,000+ humanoids deployed. over 600 industrial sites running their quadrupeds. $140M+ revenue, profitable every year since 2020. for perspective: no Western humanoid competitor is profitable. not one. OK. now here's where the "most concerning" part of this starts... if you watched the DJI story unfold, you already recognize the shape. affordable Chinese hardware quietly saturates global markets. years later, the national security questions arrive, after the install base is already massive. drones, then EVs, then AI. now robots. Unitree is running this exact playbook in real time. in April 2025, researchers found an undocumented backdoor in their Go1 robot. a remote tunnel letting anyone control the robot and stream its camera feed. default password: pi/123. 1,919 vulnerable units exposed globally. including machines at MIT, Princeton, and Carnegie Mellon. but it gets worse. every Unitree robot shares the same hardcoded encryption key. encrypt the word "unitree" and you get root access to any of them. one compromised robot can spread to every Unitree robot in Bluetooth range automatically. a literal robot botnet. the G1 quietly transmits sensor data to Chinese servers every five minutes. audio, video, GPS, LiDAR spatial mapping, with no notification, no consent, no opt-out. PLA footage has shown Go2 robots with mounted weapons. Ukrainian forces literally deployed weaponized units on the actual frontline. and every member of the bipartisan House China Committee signed a letter calling for Unitree's military company designation. Wang signed a 2022 pledge alongside Boston Dynamics not to weaponize robots. but pledges don't survive contact with shipping hardware to open markets. and under China's 2025 rules restricting military-related speech, Unitree couldn't publicly confirm PLA use even if they wanted to. 50,000+ of these robots are now deployed globally. some at institutions that probably should've asked harder questions before connecting them to their networks. the security stuff is real and people should know about it. but i also think it's important not to let that overshadow what's actually been built here. a 35-year-old who failed his English exam created a robotics company that's outshipping and outpricing every Western competitor while being the only profitable humanoid maker on Earth. most impressive and most concerning company in the world right now.

Ole Lehmann

122,620 görüntüleme • 6 ay önce