Announcing Habitat 3.0, simulating humanoid avatars and robots collaborating!... - Humanoid sim: diverse skinned avatars - Human-in-the-loop control: mouse/keyboard or VR - Tasks: social navigation and rearrangement Over 1,000 steps per second on 1 GPU for large-scale learning!show more

Dhruv Batra
261,514 次观看 • 2 年前
BOOM! Humanoid Robots Just Performed Surgery for the First... Time! REAL VIDEO! In a groundbreaking preclinical breakthrough, researchers at UC San Diego have achieved what many thought was years away: teleoperated humanoid robots successfully completing live surgeries. Published in Nature, the study marks the world’s first use of humanoid robots for in-vivo laparoscopic procedures on large animals (pigs). Two separate surgeries were completed: Key Details •. Procedure: Laparoscopic gallbladder removal (cholecystectomy) •. Team 1: Human surgeon + one humanoid robot (the robot performed core tasks while the human assisted) •. Team 2: Two humanoid robots working together with no human at the operating table •. Robots: Custom “Surgie” humanoids (~5 ft tall, ~60 lbs) using standard surgical tools •. Control: Fully teleoperated by surgeons (remote human control, not autonomous) •. Significance: First demonstration of humanoid robots handling real surgical workflows in a live setting, proving compatibility with existing OR tools and spaces This proof shows humanoid robots could one day help address surgeon shortages, enable remote procedures in rural areas, battlefields, or even space all at a fraction of the cost and space of traditional surgical robots like da Vinci. Read the full publication here: Project page with video: The future of surgery just got a whole lot more interesting. And medical cost for the first time in decades will be scheduled to go down, much further down.show more

Brian Roemmele
107,600 次观看 • 2 个月前
Announcing our commercial partnership with Booster Robotics Booster builds... humanoid robot hardware, OS, and developer tools to make humanoid robots more affordable, reliable, and practical. The partnership centers on using simulation to multiply the value of real-robot data—expanding teleoperation demonstrations into scalable training data across diverse tasks and scenes. This joint effort powers sim-real co-training and foundation-model development, accelerating progress from hardware iteration to deployable robot policies. Together, we're building simulation‑powered data infrastructure for Physical AI — making scalable training data accessible to model developers and the broader robotics ecosystem.show more

Axis Robotics
69,167 次观看 • 1 个月前
China now has its own “Bolt” — a robot... named after sprint legend Usain Bolt. A Chinese research team has unveiled the world’s first full-size humanoid robot to reach a peak speed of 10 meters per second, setting a new global benchmark for humanoid running. Bolt runs like a body pushed to the limit. Its joints and power systems work in tight coordination, keeping it balanced even at sprint speed. Built to match the build of an adult man—1.75 meters tall and 75 kilograms—it is a life-sized system operating at the edge of physics. Compared with Usain Bolt’s iconic 9.58-second 100-meter world record, which many experts believe may stand for decades, the gap between humans and machines is narrowing fast. Chinese robots are now challenging the ceiling of human performance—much as AlphaGo once challenged Go champion Ke Jie. The breakthrough builds on earlier world-record achievements in high-speed robotic running and marks a giant leap for China in humanoid motion and control. Beyond records, Bolt also carries practical value: robots are leaving the lab and stepping into real-world settings—sports training, emergency response, and demanding industrial tasks where speed, balance and control truly matter.show more

Sinical
111,438 次观看 • 7 个月前
System ID for legged robots is hard: (1) Discontinuous... dynamics and (2) many parameters to identify and hard to "excite" them. SPI-Active is a general tool for legged robot system ID. Key ideas: (1) massively parallel sampling-based optimization, (2) structured parameter space, and (3) active exploration based on Fisher Information to collect the most informative data in real. SPI-Active provides an accurate robot model and effectively reduces the sim2real gap. In sim2real policy learning setting, it outperforms baselines by 42-63% in various quadruped & humanoid tasks. Led by Nikhil Sobanbabu Guanqi Heshow more

Guanya Shi
21,442 次观看 • 1 年前
It's been incredible to see neural networks working so... well on our humanoid robots Humanoids are crazy complex - an individual motor can rotate 360 degrees and you have 40+ joints. If you do the math, that means more possible robot states than atoms in the universe Figure has our own AI model called Helix that we've designed in-house. A single Helix neural network now outputs both manipulation and navigation, end-to-end from language and pixel input Every leap in machine learning has come from massive, diverse datasets. At Figure, we’re currently building the largest pretraining dataset for humanoids in history - excited to see what this unlocksshow more

Brett Adcock
93,986 次观看 • 11 个月前
📢Announcing our 3D head avatar benchmark📢 Two tasks with... hidden test sets: - Dynamic Novel View Synthesis on Heads - Monocular FLAME-driven Head Avatar Reconstruction Our goal is to make research on 3D head avatars more comparable and ultimately increase the realism of digital humans. The benchmark studies distinct phenomena of 3D head avatar creation, such as extreme facial expressions, slow motion captures of shaking long hair, or complicated light reflection and refraction patterns of glasses. The two benchmark tasks assess two core desiderata of 3D avatars: While the novel view synthesis challenge focuses on best possible rendering quality of complex moving scenes, the avatar animation challenge is concerned with how well a driving signal is translated into an avatar. Evaluations are light-weight and consist of diverse video recordings from the popular NeRSemble dataset with a hidden test set. Participation in the benchmark is therefore straight-forward and requires only 5 reconstructions per task. Leaderboard and benchmark submission: Benchmark data access and toolkit: Great work by Tobias Kirschstein Simon Giebenhainshow more

Matthias Niessner
28,105 次观看 • 1 年前
Genesis AI just unveiled Eno. It's humanoid robot that... challenges everything the industry assumed about what robots should look like. Forbes just called it 'the iPhone moment for humanoid robots'. No head. No face. No exposed motors or cables. 22 degrees of freedom per hand with different finger lengths (like actual human hands). Back-drivable for safety. Onboard cameras and tactile sensors. In demos: bundling wires with tape (genuinely hard, tape is sticky and unpredictable), performing lab automation with millimeter precision on unmodified equipment. Optional chest screen shows the robot's reasoning before it acts, a visual window into its mind to build trust. Powered by Genesis AI GENE foundation model. Payload 3-5kg per arm, 4-6 hours battery. Industrial deployments late 2026, homes much later. ~~ ♻️ Join the weekly robotics newsletter, and never miss any news →show more

Lukas Ziegler
29,453 次观看 • 2 个月前
home page hero ✨ Design notes: - "Forever" hero... text dot pixel FX done in Unicorn Studio. (I will do a whole tutorial on this later. Unicorn's WebGL engine is absolutely wild and very powerful / robust) - built in Framer - I wanted to recreate the iOS unlock effect where your home screen icons cascade into place in a beautifully timed choreography. This took a lot of careful timing using Framer's "Appear" effect on the hero text and surrounding avatars because it was super important that we didn't lose the legibility of our main message ("Build Your Forever Audience") with all the animations. - If you look closely, the choreography is setup to lead your eye through the hero text first starting with "Build Your" then "Forever" and finally "Audience." - With those text layers in place + the surrounding avatars, there is a slight 1 sec pause before the remaining elements slide in below and above (How it works, CTA buttons, announcement badge, and lastly the main nav). - All told the entire loading sequence is 6 seconds - Custom particle system powers the interactive star field (the stars slowly gravitate to your pointer position, and the star field perspective changes ever so subtly as you move your mouse around on the page) - I have 3 shooting stars made of small white line layers that start out off canvas rotated at different angles that shoot across to another point off canvas at random times on a loop effect. - Given this hero scene is in space, I wanted the surrounding avatar elements to "float" in low gravity mode. For this I used Framer's loop effect that slowly oscillates the layer's y position. I then offset the delay of each element randomly to stagger the floating loop so each avatar floats independently/randomly - The final major treatment for this hero scene was the scroll animations on the avatars. I wanted to create a bit of a warp speed effect when you scroll down, as if the avatars were being pulled or sucked into a worm hole as you scroll down below this hero fold. - To accomplish this, I applied Framer's scroll transform affect set to "section in view" on each of the floating avatars, and set the "scroll to" position of the upper avatars to be much, much further away on the y-axis than the "scroll to" position of the lower avatars. (eg. -1700px on upper most avatars vs. -600px on lowest positioned avatars). This effectively causes the upper avatars to slide up off the hero canvas with much greater velocity than the lower positioned avatars. - And when you scroll back up to the hero section, the inverse happens where the lower positioned avatars "arrive back in place" from up above the hero canvas before the upper avatars come back into the scene and settle in place. - Overall I wanted this hero section to feel alive. The floating avatars, particle system with very subtle star movements, and the Caustics effect on the "Forever" text all sort of move at the pace of slow breathing - which is a great pace to create a sense of life and comfort in your scene. Conversion Results (so far) - When this new Framer site launched along with Calaxy v1.9 release on Base a couple weeks ago, we saw a surge in traffic, around 20k page views in the first few days. - Of those 20k hits, 11k visited the app install page ( - which is our main CTA - We saw around 10K new users in the first week after v1.9 launch Overall I'm very happy with the new site and early performance metrics. Lots of tweaking to do but its a good start. If you are a designer building in Framer - hit me with any questions on the above hero notes. Happy to share more specifics! 👾show more

Chadd Weston
16,365 次观看 • 10 个月前
TESLA HALTED MODEL S AND MODEL X PRODUCTION TO... BUILD AN ARMY OF OPTIMUS ROBOTS The Fremont assembly line was torn down in 46 days. In its place, Tesla is building a line for humanoid production, aiming for a million units a year A humanoid robot is a body shaped like a human. Physical AI is the intelligence that controls that body Walking and making coffee is often just imitation learning from a scripted routine. But once the environment shifts, the learned trick stops working Language models had the entire internet to train on. Robotics has nothing close to that scale of data, which is why one giant brain hasn't worked for anyone yet The industry is moving toward modularity instead - separate models for vision, movement, and planning, each improved on its own The real question is no longer whether a robot can move impressively. It's whether it can pull its sensors into one picture of the world and adapt to whatever wasn't scripted for itshow more

iamigorekk
22,205 次观看 • 24 天前
🦿Xpeng showed a humanoid robot called IRON whose movement... looked so human that the team literally cut it open on stage to prove it is a machine. IRON uses a bionic body with a flexible spine, synthetic muscles, and soft skin so joints and torso can twist smoothly like a person. The system has 82 degrees of freedom in total with 22 in each hand for fine finger control. Compute runs on 3 custom AI chips rated at 2,250 TOPS (Tera Operations Per Second), which is far above typical laptop neural accelerators, so it can handle vision and motion planning on the robot. The AI stack focuses on turning camera input directly into body movement without routing through text, which reduces lag and makes the gait look natural. Xpeng staged the cut-open demo at AI Day in Guangzhou this week, addressing rumors that a performer was inside by exposing internal actuators, wiring, and cooling. Company materials also mention a large physical-world model and a multi-brain control setup for dialogue, perception, and locomotion, hinting at a path from stage demos to service work. Production is targeted for 2026, so near-term tasks will be limited, but the hardware shows a serious step toward human-scale manipulation.show more

Rohan Paul
3,802,543 次观看 • 10 个月前
ENGINEAI just opened registration for URKL, a global humanoid... fighting league with an insane ¥10,000,000 (approx. $1.39 million) top prize. 🤖🥊 This is a massive engineering challenge focused on motion control and balance using the "T800" humanoid as the standard bot. The rules are strictly "non-violent," meaning no destructive mods are allowed. You win through better code and smarter protective gear. Here is the breakdown for teams looking to jump in: ➤ Massive Payouts: The winner takes ¥10,000,000 (approx. $1.39 million), second gets ¥2,000,000 (approx. $278,000), and third takes ¥1,000,000 (approx. $139,000). ➤ Hardware Perks: Every team that makes it into the Top 16 officially owns their T800 robot. ➤ Career Fast-Track: Top 8 finalists get a "Green Channel" straight to the final interview for job offers at ENGINEAI. ➤ Registration: Open from March 1 to April 30. Teams need at least 3 members with skills in control, electronics, or mechanical design. ➤ Global Finals: After the qualifiers, the world championship is set for December 2026 through January 2027. Once you are in, the committee hands over the simulation platform and T800 models to start training your boxing algorithms. Full Info: #Robot #Humanoid #Robotics #AI #EmbodiedAI #PhysicalAI #URKL #ENGINEAI #RobotFightingshow more

RoboHub🤖
30,637 次观看 • 6 个月前
This work makes a humanoid robot do simple parkour... moves by looking with a depth camera and choosing the right move on the fly. The big deal is that it turns lots of small human moves into long, real-time robot behavior, without hand-coding every transition or retraining for each new course. A humanoid robot is usually good at steady walking, but it often fails when it has to do fast moves like jumping up, vaulting, or rolling, and then keep going to the next obstacle. The hard part is that you cannot easily collect training data for every possible obstacle shape, distance, and mistake, so robots end up learning a few moves that only work in a narrow setup. This work starts from short clips of real human parkour moves, like stepping over, vaulting, climbing, and rolling. It uses motion matching, which is basically a smart “pick the next clip that fits best right now” search, to stitch those short clips into a long, smooth plan that looks like a human doing a whole course. Then it trains a controller with reinforcement learning (RL), which means the robot learns by trial and error to copy that plan while staying balanced and not falling. After training separate expert controllers for different moves, it compresses them into 1 controller that uses only onboard depth sensing and a simple “go this fast in this direction” command. In real tests on a Unitree G1 humanoid, it can clear multiple obstacles in a row, adapt when obstacles get moved, and climb a wall up to 1.25m.show more

Rohan Paul
37,121 次观看 • 6 个月前
Robotics keeps hitting the same wall. Single task RL... works, but... it does not scale to hundreds of tasks or new embodiments. This new paper looks like a real step toward fixing that. The team introduces MMBench, a benchmark with 200 tasks across many domains and robots, and Newt, a language conditioned world model trained online across all 200 tasks at once. The simple idea behind Newt: The model learns from demos to get the right priors It trains across many tasks through online interaction It uses language to ground the goal It adapts fast when a new task shows up What stood out to me: ✅ One model trained on 200 tasks at the same time ✅ Language conditioned control for both states and RGB ✅ Better data efficiency than strong baselines ✅ Strong open loop control ✅ Fast adaptation to new tasks and embodiments ✅ Full release of 200 checkpoints, 4000 demos, code, and benchmark This is a good push toward general control instead of one model per task. If you want the full paper: Project page: —- Weekly robotics and AI insights. Subscribe free:show more

Ilir Aliu
70,090 次观看 • 9 个月前
Does LLM really need to be a helpful assistant... all the time? No. If you want to simulate people, “perfectly helpful” could be the wrong objective. Meet OdysSim, a journey toward LLMs beyond assistants, as behavioral foundation models (10B tokens of real human behavior; 23 sim benchmarks, finally in one place. new open models: outperform or on par with GPT-5.5, Gemini 3.1, or Claude Opus 4.7 in many behavior-sim dimensions). Human behavior simulation is becoming essential. Agent evaluation needs realistic users before real users show up. Medical and classroom training need realistic patients and students. Social science needs synthetic participants at scale. But real people are not ideal assistants. Real patients panic or ignore good advice. Real students misunderstand. Real customers are vague, picky, impatient, or simply leave. Human behavior is messy, diverse, and often imperfect. Frontier LLMs are getting better at math, code, and long-horizon tasks. They are NOT getting better at simulating human behavior. If anything, they drift the other way: more assistant-ish, more homogeneous, fewer of the errors and quirks real humans show. This is no accident. The whole pipeline is built for helpfulness and task success, not behavioral realism. And you can't prompt your way out of that. So we rethink the recipe from scratch and release: 🧠 The OdysSim corpus: 21.4M real human interactions (~10B tokens) from 62 sources, every conversation retrofitted with social grounding (who is talking, and why) 📏 SOUL-Index: 23 human-behavior benchmarks unified into one suite across 5 axes 🤖 OSim-8B: open weights; tops more SOUL-Index benchmarks than any frontier model, acts more like a real user than any of them on τ-bench (nearly matching real humans in the reaction dimension), and writes far more human-like text along the way.show more

Xuhui Zhou
142,830 次观看 • 3 个月前
🚀 Introducing EgoExo Forge - built on top of... Rerun, Gradio, and Hugging Face hub (I’ll be in San Francisco July 21–29 — if you’re into robotics, egocentric AI, large-scale data collection, or just want to chat, DM me!) In my opinion, large-scale, diverse, and high-quality data is still the largest bottleneck for generalized robotics deployment. I believe that some version of imitation learning from human examples will be the most scalable + clean way to train humanoid robots 🤖 (similar to what Tesla did for Full Self Driving). Teleop is too expensive to collect a large enough dataset in a reasonable manner, so passive collection via egocentric (and in certain cases, exocentric) views feels like the right bet. Over the past few months, I've been trying to build out the scaffolding for this and using Rerun as my underlying infrastructure. Data being collected needs to be easily inspectable + time series and rerun provides the right tooling for this. My goal is to first build out a ground truth representative dataset from already existing open source data, generate some reasonable baselines, and then go out and collect my own data that adheres to the defined schema. 🔍 Starting with open-source datasets 1. EgoDex from Apple 2. HOCap from Nvidia and the University of Texas at Dallas 3. Assembly101 from Meta All these different datasets have different sensor configurations + annotations, so my goal with egoexo-forge is to have one consistent labeling scheme + data layout. I built a data pipeline that aligns all of the different datasets in one general schema assuming the COCO133 keypoint layout that allows for exo+ego, ego only, or exo only Since the scaffolding is already there, it becomes MUCH easier to add other datasets. So the next ones that I'll be including are HD-EPIC kitchens dataset, HOT3D, and finally my own personal iPhone + insta360 go collection method. Once I have a diverse variety of datasets, I'll double down on what I believe to be the key algorithms required to make useful data for imitation learning 📊 1. Camera Pose estimation via SLAM/SFM for ego perspective (and automatic calibration for exo) 2. Human pose estimation for both egocentric + exocentric views 3. Metric 3D reconstruction + object tracking I'll be setting up reasonable open-source baselines for each of these to validate that these datasets work, and then finally try to use the generated datasets for some imitation learning via the pi0-lerobot repo I've been working on. I plan on making a blog post + providing more info on all of this in the near future so stay tunedshow more

Pablo Vela
35,919 次观看 • 1 年前
Imagine having a ping pong robot! 🏓 Researchers and... developers building physical AI: meet Reachy 2 from Pollen Robotics, an open-source, humanoid robot for real-world experimentation. It’s a bimanual mobile manipulator: each 7-DOF arm mimics human proportions and can lift up to 3 kg, giving dexterity for object handling. It can be controlled with Python and ROS2 Humble, or go straight into VR teleoperation, use a headset to move Reachy’s arms, hands, and head, and see through its cameras as if you’re in the robot’s own body. Want it to move around? A mobile base with three omnidirectional wheels, rich sensors, and LiDAR lets Reachy 2 navigate and explore its surroundings smoothly. 🗺️ Under the hood, it’s powered by a CPU system that’s ready for machine learning, perfect for loading AI frameworks and testing new models from Hugging Face directly on the robot. Keep making robots more, and more accessible Pollen team! ... and keep making more open source models to make robots more mainstream clem 🤗!show more

Lukas Ziegler
37,221 次观看 • 1 年前
Microsoft presents Windows Agent Arena Evaluating Multi-Modal OS Agents... at Scale discuss: Large language models (LLMs) show remarkable potential to act as computer agents, enhancing human productivity and software accessibility in multi-modal tasks that require planning and reasoning. However, measuring agent performance in realistic environments remains a challenge since: (i) most benchmarks are limited to specific modalities or domains (e.g. text-only, web navigation, Q&A, coding) and (ii) full benchmark evaluations are slow (on order of magnitude of days) given the multi-step sequential nature of tasks. To address these challenges, we introduce the Windows Agent Arena: a reproducible, general environment focusing exclusively on the Windows operating system (OS) where agents can operate freely within a real Windows OS and use the same wide range of applications, tools, and web browsers available to human users when solving tasks. We adapt the OSWorld framework (Xie et al., 2024) to create 150+ diverse Windows tasks across representative domains that require agent abilities in planning, screen understanding, and tool usage. Our benchmark is scalable and can be seamlessly parallelized in Azure for a full benchmark evaluation in as little as 20 minutes. To demonstrate Windows Agent Arena's capabilities, we also introduce a new multi-modal agent, Navi. Our agent achieves a success rate of 19.5% in the Windows domain, compared to 74.5% performance of an unassisted human. Navi also demonstrates strong performance on another popular web-based benchmark, Mind2Web. We offer extensive quantitative and qualitative analysis of Navi's performance, and provide insights into the opportunities for future research in agent development and data generation using Windows Agent Arena.show more

AK
19,684 次观看 • 2 年前
𝗥𝗼𝗯𝗼𝘁𝘀 𝗱𝗼𝗻’𝘁 𝗻𝗲𝗲𝗱 𝗺𝗼𝗿𝗲 𝗱𝗲𝗺𝗼𝗻𝘀𝘁𝗿𝗮𝘁𝗶𝗼𝗻𝘀. 𝗧𝗵𝗲𝘆 𝗻𝗲𝗲𝗱 𝘁𝗼 𝗹𝗲𝗮𝗿𝗻... 𝗳𝗿𝗼𝗺 𝗳𝗮𝗶𝗹𝘂𝗿𝗲 — 𝗮𝗳𝘁𝗲𝗿 𝘄𝗮𝘁𝗰𝗵𝗶𝗻𝗴 𝗵𝘂𝗺𝗮𝗻𝘀. Most robot learning systems assume failure is the end of learning. In our new work, we study whether robots can improve after deployment by learning from their own failures, without any human intervention, teleoperation, or corrective labels. The key idea is simple: human videos contain structure about how the world works. We use them to learn cross-embodiment representations of action, dynamics, and value, enabling a shared predictive space between human behavior and robot experience. This allows a new learning loop: 👉 pretrain on human videos 👉 deploy robot policy 👉 observe failures 👉 reinterpret failures using human priors 👉 improve autonomously We evaluate this across 7 real-world manipulation tasks, showing: 📈 40% → 81% success rate 🏆 Strong improvements over π0.6 RECAP and RISE ✔️ Zero human intervention during post-deployment improvement 🧬 Generalizes across robot embodiments and policy backbones A key finding is that explicit failure repair significantly outperforms failure reweighting, yielding substantially larger gains under identical data conditions (+25 pts vs +5 pts on the same π0.5 base policy). Overall, the results suggest a shift in how we think about robot learning: Human videos are not only for pretraining policies. They can provide the structure needed for continual self-improvement after deployment. 📄 Paper: 🌐 Project: I am grateful for working with the fantastic leads Hanzhi Chen and Anran Zhang, and our collaborators Simon Schaefer, Kejia Chen, Shi Chen, Daniel Cremers. Special thanks to Stefan Leutenegger for co-advising this project with me. ETH Zürich TU München Microsoft Check out Hanzhi's 🧵 for more detailsshow more

Oier Mees
12,514 次观看 • 2 个月前
This is where humanoid robots start becoming economically interesting.... 🤖🏭 Not in a demo stage. Not doing backflips. Doing repetitive industrial work beside humans. The real breakthrough for humanoid robotics will come when machines can reliably handle tools, materials and production tasks designed for human workers — without factories needing to rebuild everything around them. That transition could redefine manufacturing. Would you trust humanoid robots on a production line? 👇 #HumanoidRobots #Robotics #AI #Automation #5G #IoT #NLP #Tech #Technology #innovation #Bigdata Cc: Jean-Baptiste Lefevre Nicolas Babin Pinna Pierre Rosy Eric Gaubert Yann Marchand Eveline Ruehlin Franco Ronconi 🇮🇹 Dev Khanna Dr. Khulood Almani | د.خلود المانع Dr. Marcell Vollmer #StaySafe #CES2026 Joanne Moretti. Margaret🌴Siegien 🐦📷 Shi💙 Anand Narang Elitsa Krumova Harold Sinnott 📱 BusinessIntelligence Evan Kirstel #B2B #TechFluencer Hana Sen. Sally Eaves Spiros Margaris ipfconline Laurent Alaus Mack Jeff KAGAN Industry Analyst, Strategic Advisor Andres Vilariño 🇪🇦 Dr Efi Pylarinou Françoise Morvan MHcommunicate #Mastodon👉@mhcommunicate@social JC Gaillard Archon Security Jola Burnett Jérôme MONANGE Mary Gambara Dr Stephen Harwood #Tech4Good #SDG 🇵🇱#CES2026 Elinor Stutz Fati Sule Dr. Debashis Dutta Dinis Guarda Devaang Bhatt Terence Mills @theomitsa Tony Moroney #DigitalTransformationshow more

Fabrizio Bustamante Escudero
43,973 次观看 • 1 个月前