Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Introducing the QUANXTA Zero Series, X Square Robot’s next-generation UMI data collection solution for embodied AI. The series includes three products: QUANXTA Zero-G0: VR headset + backpack + dual grippers QUANXTA Zero-G1: headband rig + dual grippers QUANXTA Zero-E0: headband rig Built for scalable embodied AI data production: Multimodal...

354,518 Aufrufe • vor 1 Monat •via X (Twitter)

0 Kommentare

Keine Kommentare verfügbar

Kommentare vom Original-Post werden hier angezeigt

Ähnliche Videos

🔥 JUST IN: Open-source robotics dataset from 100% real-world scenarios! 🤯 Chinese robotics company AGIBOT just released AGIBOT WORLD 2026, an open-source dataset systematically covering key embodied AI research directions. Built entirely from real-world environments: commercial spaces, and homes. Collected using AGIBOT G2 robots in free-form collection mode, providing structured, accurately annotated, high-quality data. Digital twin technology creates 1:1 scale replicas in simulation matching the real environments. Both real-world and simulation data are open-sourced. The AGIBOT G2 platform collects multiple data types simultaneously: RGB(D) cameras, tactile sensors, force sensors, LiDAR, IMU, and full-body joint states. Whole-body control coordinates arms, waist, and hands for complex tasks. First-person teleoperation lets operators control the robot from its perspective. The tasks covered are fine-grained manipulation, ultra-long-horizon tasks, spatial navigation, dual-arm coordination, and multi-agent/human-robot collaboration. The dataset includes error-recovery trajectories with annotations. Most datasets only show successful demonstrations. AGIBOT includes failures and how the robot recovers, teaching models how to handle mistakes. After collection, data is tested through policy training and real-robot deployment to ensure quality. Then processed through industrial quality control with multiple screening and cleaning rounds. Making it open-source accelerates embodied AI research by giving researchers access to high-quality real-world robot data at scale. 🇨🇳 Learn more here: ~~ ♻️ Join the weekly robotics newsletter, and never miss any news →

Lukas Ziegler

40,583 Aufrufe • vor 4 Monaten

🚀 We just raised $40 million to build infrastructure for Physical AI! 🦾 AI is rapidly transforming critical industries like manufacturing, logistics, transportation, agriculture, construction, aerospace, and defense. Teams that win in the physical world are those who can create a data flywheel, leveraging infrastructure to capture, ingest, analyze, and evaluate the vast quantities of data generated by real-world systems. Robotics data is multimodal, time-synchronized, and bandwidth‑constrained at the edge. Traditional data and observability platforms were only designed to store and query text and time-series data, not petabyte-scale 3D, video, audio, GNSS, and proprioceptive data. The ability to efficiently capture, ingest, search, visualize, and evaluate multimodal data is critical to Physical AI development. Foxglove is a modern data engine for Physical AI, enabling you to record logs or capture demonstrations at the edge, sync recordings to the cloud or on-premises storage, find critical events across petabytes of data, evaluate robot performance, and watch a 3D frame-by-frame replay using our advanced visualization tool. 👉 Today is still Day 1 for Physical AI, and we're hiring for dozens of roles to assemble the best team in the industry. If you've built ML platforms, data infrastructure, dataset curation, evaluation and validation, or visualization tools at a leading robotics or autonomous vehicle company, let's chat – drop me a note or tag a friend below and I'll follow up personally! Thank you to Alexandra Sukin and Jeremy Levine at Bessemer, Seth Winterroth 🤖 at Eclipse, David Beyer and Sunil Dhaliwal at Amplify Partners, and Icehouse Ventures for joining us on this mission. Also a special shoutout to our angels tobi lutke Alex Kendall Kyle Vogt Milan Kovac Hussein Mehanna Pieter Abbeel Brad Porter Boris Sofman Kevin Peterson Chris Walti Lindon Gao Daniel Kan Adam Draper ⏻ Fred Ehrsam and Karri Saarinen!

Adrian Macneil — 🤖/acc

46,584 Aufrufe • vor 9 Monaten

X Square Robot just closed its Series C at a valuation above RMB 20 billion, about $2.8 billion 🤖 IDG came into this round. The bigger signal is the cap table. HongShan and Xiaomi were already in across earlier rounds, while Meituan, Alibaba, ByteDance, and Xiaomi have each led rounds at different stages. That puts X Square in a rare position for an embodied AI company: top-tier financial capital on one side, and four of China’s biggest tech platforms on the other. This is not just a money story. Meituan, Alibaba, ByteDance, and Xiaomi bring very different strategic assets: real-world scenarios, cloud infrastructure, consumer traffic, supply chains, and hardware ecosystems. The deployment side is already moving: robot home-cleaning services first, then a “Robots Into Homes” program with the first batch entering real households. The model stack is worth watching too. X Square has open-sourced WALL-OSS-0.5 for robot manipulation and WALL-WM for world modeling. WALL-OSS-0.5 showed strong real-robot performance without post-training, while WALL-WM uses event-level prediction to align language, vision, and action around meaningful physical-world events. They are also building a model-driven data pipeline for large-scale collection, cleaning, annotation, quality control, and augmentation. That matters because home robotics dies in the long tail: weird rooms, messy objects, bad lighting, and tasks that never look the same twice. Founded in 2023, X Square is building general-purpose embodied AI robots and foundation models for real-world environments, tying models, robot hardware, high-precision manipulation, data, and deployment into one system.

RoboHub🤖

12,975 Aufrufe • vor 2 Monaten

We believe we’re the first robotics company to demonstrate a robot peeling an apple with dual dexterous human-like hands. This breakthrough closes a key gap in robotics, achieving bimanual, contact-rich manipulation and moving far beyond the limits of simple grippers. 🧵↓ Today’s AI models (VLMs) are excellent at perception but struggle with action. Controlling high-degree-of-freedom hands for tasks like this is incredibly complex, and precise finger-level teleoperation is nearly impossible for humans. Our first step was a shared-autonomy system: rather than controlling every finger, the operator triggers pre-learned skills like a “rotate apple or tennis ball” primitive via a keyboard press or pedal. This makes scalable data collection and RL training possible. How does the AI manage this? We created "MoDE-VLA" (Mixture of Dexterous Experts). It fuses vision, language, force, and touch data by using a team of specialist "experts," making control in high-dimensional spaces stable and effective. The combination of these two innovations allows for seamless, contact-rich manipulation. The human provides high-level guidance, and the robot executes the complex in-hand coordination required. This work paves the way for robots that can safely handle delicate tasks in human environments. Want the full technical details? 📄 Read the full research paper: Visit us at NVIDIA GTC Booth #1838, Hall 3 to learn more! #Robotics #AI #DexterousManipulation #VLA #NVIDIAGTC Nancy Villicaña NVIDIA GTC

Sharpa

20,429 Aufrufe • vor 5 Monaten

Learning from Human Demonstrations: Show the Robot How to Act! The pipeline is very similar to older experiments using Gemini & pi0 with LeRobot. Pi-zero runs locally, while Gemini Flash generates the affordances and the high-level task. (More details are in the thread.) The new component is learning from demonstrations via Gemini 2.5 Pro. I capture a video while demoing & take one of the last frames. Gemini 2.5 Pro then extracts the instructions & passes them to Gemini Flash to process the scene. The fun part is that there's no fancy insight that came from me; other than the days spent figuring out the right prompts. It's the bitter lesson hitting you in the face -> Enhanced Gemini capabilities make this possible. For example, Gemini Flash cannot do Russian doll stacking, but Gemini 2.5 Pro can do it consistently. The current limitation is low-level manipulation: - As you can see, I'm aligning the objects so they are easy to grasp using the same technique from the training data. I couldn't get Gemini Flash to consistently output an accurate grasping angle, and Gemini 1.5 Pro was too expensive and slow for real-time deployment. - Getting a symmetrical gripper should also help a lot. Adding rubber to the tips would probably also help prevent objects from slipping. Collecting & curating the data was the most time consuming & labor intensive part. Next, to improve low-level manipulation and make the system more real-time, I'm shifting to focus more on sims & synthetic data. This aligns better with my core competence. I'm open to tips and suggestions.

Shreyas Gite

22,555 Aufrufe • vor 1 Jahr

X Square Robot Unveils New Embodied AI Model, Says Robots Will Arrive in Homes in 35 Days Backed by Alibaba, ByteDance, Xiaomi and Meituan, X Square Robot unveiled a next-generation embodied AI foundation model for home robots and said its first deployments in everyday households will begin within 35 days. X Square Robot on Tuesday unveiled WALL-B, a new embodied AI foundation model designed for deployment in real-world homes, marking what the company described as a major step toward bringing general-purpose robots into daily family life. At a launch event themed "Born to Bot, Bot to Family," the company also introduced its World Unified Model (WUM) architecture, a training framework that combines vision, language, action and physical prediction within a single system from the outset. X Square said the model is intended to help robots operate in the far more unpredictable setting of a home, where tasks, layouts and interactions vary from moment to moment. "Robots in factories and in homes are completely different. In factories, they repeat the same action 10,000 times without variation. In a home, however, they need to perform 10,000 different actions, each unique and non-repetitive. Therefore, the challenge of a truly intelligent robot lies not in repeating a single action, but in the ability to execute new, untrained movements within unstructured environments. Deploying robots in the home is one of the most significant technical hurdles of our time," said Qian Wang, founder and CEO of X Square Robot. WALL-B is the first real-world implementation of the World Unified Model architecture. Unlike modular systems that train perception, language and control separately, X Square Robot said World Unified Model optimizes those capabilities jointly from the very beginning. The company said that allows physical prediction — including force, friction and collision dynamics — to emerge as part of the model itself, rather than being layered on afterward. "We train all capabilities—vision, language, action, and prediction—within the same network from day one. Much like infants, who do not learn to see, move and speak in isolated, sequential stages, but instead see, move listen and act simultaneously while receiving feedback, we have integrated all these capabilities into a unified whole," said Wang Hao, CTO of X Square. X Square Robot said the development of WALL-B rests on two pillars. The first is a data strategy that prioritizes training on authentic, non-staged home environments to cover the “long-tail” distribution of real-world scenarios, such as misplaced objects and temporary occlusions. Unlike models primarily trained on synthetic data or laboratory datasets, this strategy exposes WALL-B to the natural clutter of lived-in spaces—misplaced items, unexpected obstacles, and spontaneous human activity—ensuring that the training data reflects real-world conditions rather than a simplified version. The second is a physics-aware predictive mechanism that anticipates physical outcomes before an action is taken, enabling the model to respond to contact dynamics instead of just reacting. The development of the self-developed WUM architecture on physical robotic platforms highlights the company’s accumlated experience in bridging sim-to-real gaps across varied operational contexts. Wang commented that the current AI model is still in an "intern" stage, subject to errors requiring remote assistance. For instance, it may mistakenly place slippers in the kitchen or pause while wiping a table to "think". However, the model operates nonstop 24 hours a day, becoming increasingly "intelligent" as each day of operation generates new data. In 35 days, on May 25, X Square Robot will officially bring its robots into everyday homes, underscoring the company’s long-term commitment to the home robotics sector.

X Square Robot

52,968 Aufrufe • vor 4 Monaten

China unveils humanoid robot with lifelike skin and blinking eyes built for daily life | Prabhat Ranjan Mishra, Interesting Engineering Large Language Models (LLMs) and Vision-Language Models (VLMs) help process and interpret complex data from human interactions. A Shanghai-based company has developed humanoid robots that appear as real as humans. The advanced bionic humanoid robot is integrated with self-supervised AI algorithms. Named Elf V1, the robot can perceive the world, communicate, learn, and interact intelligently with its surroundings. Developed by AheadForm Technology, the robot offers up to 30 degrees of freedom, powered by a precise control system and an advanced AI learning algorithm. Robot offers expressive facial features The robot offers expressive facial features, moving eyes, and synchronized speech. It can also convey emotions and understand human non-verbal cues, making interactions more natural and engaging. The robot has highly interactive capabilities and lifelike appearances. AheadForm expects that its robots could soon seamlessly integrate into daily life, providing assistance, companionship, and support across various industries. “We believe that by developing realistic and expressive robot heads, we can bridge the gap between humans and machines, fostering a new era of interactive and intelligent robotics,” said the company in a statement. Reports revealed that to avoid the “uncanny valley” effect and be able to interact with us, they are given lifelike skin and capabilities to read our emotions and respond appropriately using dynamic expression simulation and emotion generation tech. Bionic skin and high-precision control system The Elf V1 series of humanoids features 30 facial muscles animated by brushless micro-motors and managed by a high-precision control system. Paired with an ability to detect their users’ emotions with low latency and bionic skin, their facial expressions are nearly identical to those of humans, reported CGTN. The company claims it’s pioneering the development of realistic humanoid robots designed to revolutionize human-robot interaction. It’s enhancing sophisticated humanoid robot heads that can express emotions, perceive their environment, and interact seamlessly with humans. By combining cutting-edge AI and advanced robotics, AheadForm aims to bring life to machines and transform how humans engage with technology. AI models boost robots’ responsiveness Seamless integration of Large Language Models (LLMs) and Vision-Language Models (VLMs) into the humanoid robots can help them process and interpret complex data from human interactions, enabling the robot to learn and adapt in real-time, achieving human-level understanding and responsiveness. AheadForm uses Brushless Motors that deliver ultra-quiet operation and high responsiveness, specifically designed for precision facial movements in humanoid robots. With its compact size, lightweight design, and energy efficiency, this motor is the ideal choice for next-generation robots that require precise, subtle facial control to create a truly human-like experience. Previously, the company unveiled the Lan Series that features realistic humanoid robots with soft skin and 10 degrees of freedom, offering a lifelike appearance and intuitive movements. This series is designed for cost-efficiency, for applications prioritizing mobility and manipulation.

Owen Gregorian

179,005 Aufrufe • vor 10 Monaten

Most video-action robot models are a content-creation video generator with an action module attached. LingBot-VA 2.0 from Robbyant, a video-action foundation model, throws that starting point out and trains the whole stack natively for control. And it runs closed-loop at a peak 225 Hz. It's so important because A robot cannot move responsively when its controller pauses to imagine the next few frames. LingBot-VA 2.0 predicts during execution, then corrects using each real observation. And it carries only about 13B video parameters while activating roughly 1.9B per token. Bigger robot models usually mean slower reactions, creating a direct conflict between intelligence and control. LingBot-VA 2.0 is trained from scratch for robot control rather than adapted from a video generator built for content creation. Robbyant, an embodied AI company under Ant Group, built it to learn how scenes change under actions, predict what should happen next, and turn those predictions into real-time robot movements. Most video-action systems inherit a tokenizer and video backbone trained mainly to reproduce visual appearance. LingBot-VA 2.0 rebuilds both parts around physical control. Its semantic visual-action tokenizer maps observations toward features from a frozen vision foundation model and learns compact latent actions from frame-to-frame changes using self-supervised inverse and forward dynamics. Unlabeled web video can therefore carry action-relevant training signals without robot action labels. The policy is causal from the start, so every prediction can use only past observations. Its sparse Mixture-of-Experts video backbone has about 13B total parameters, while about 1.9B are active per token, keeping the compute lower during each step. A high-level vision-language planner breaks long tasks into smaller instructions, while the low-level video-action policy handles continuous movement. Foresight Reasoning predicts future visual states while the robot is already acting, then replaces imagined states with every new real observation. Combined with few-step distillation and systems acceleration, the paper reports a peak asynchronous execution frequency of 225 Hz. The model adapts from 10–15 demonstrations, transfers across robot embodiments, and handles some new tasks zero-shot. In the paper’s own evaluations, it reaches 93.6 average on RoboTwin 2.0 and reports stronger real-world results than LingBot-VA and π0.5 across the tested tasks. 🧵 1.

Rohan Paul

11,253 Aufrufe • vor 1 Monat

Everyone in Embodied AI is talking about Vision-Language-Action (VLA) models. Almost no one is talking about the physical nightmare of collecting the data to train them. You can't scrape a kitchen table or a warehouse shelf from a web browser. To get to millions of hours of diverse, real-world manipulation data, you need hundreds of rigs. But you can't buy them. If you build them out of off-the-shelf developer kits, they weigh 15 pounds, overheat in an hour, require bulky cabling, and break the first time an operator wears them on a job site. At the Instawork Robotics Lab (IRL), we had to build our own. Meet the Instacore: a rugged, 4lb wearable egocentric data-capture engine designed specifically to survive a full shift on a standard power pack. We didn't build a flashy humanoid. We did hard, blue-collar systems engineering: 💾 THE COMPUTE — A custom MediaTek Genio carrier board that runs completely fanless, routing 5 camera streams directly to on-board storage. ⚡ THE I/O PIPELINE — We ditched USB for industrial GMSL. Thin, ultra-flexible coaxial lines route high-speed data down to the backpack and Power-over-Coax (PoC) back up, completely eliminating batteries on the wrist. ⏱️ UNIFIED SENSOR CLOCK — The MediaTek Genio SoC drives a shared master clock straight to the ISPs driving our 5 global shutter sensors, stamping metadata at the microsecond of capture to ensure zero-drift temporal alignment. 👁️ OPTIMIZED OPTICS — 95 DFOV lenses on flexible PCB ribbon modules keep the wrist cams flat to prevent snags. A 145-degree chest camera captures the macro workspace, while a 50mm baseline rectilinear stereo head pair preserves close-up 3D mapping. To build hundreds of these, we took over a warehouse in Mountain View in April, called it the Instalab, and brought in talented Pros with assembly backgrounds from Tesla, Apple, and Meta. To test the systems, we had our Pros wear active rigs while assembling more rigs. The video below shows the raw, time-synchronized Foxglove playback of that exact loop. Now, we’re shipping these units globally to scale data collection for real Pros on the job.

Ryan Hickman

26,385 Aufrufe • vor 2 Monaten

A mysterious embodied AI demo has recently sparked a lot of discussion. In the video, two robots from different manufacturers with significantly different hardware architectures, the Unitree G1 and AgiBot Yuanzheng A3, are reportedly running on the same “brain.” In a complex indoor environment, they work continuously for around 10 minutes in a single uncut take, performing a series of long-horizon tasks including cleaning windows, organizing objects, resuming interrupted tasks, and cooperating with each other. What makes it even more interesting is that when one robot cannot reach a high shelf, it attempts to use a box to solve the problem. The two robots also appear capable of cooperating based on each other’s physical capabilities. If the claimed level of autonomy and the use of the same model across different embodiments are eventually verified, I think there are three things that really deserve attention: 1. Cross-embodiment generalization. If the same foundation model can operate two substantially different robotic platforms, it could mean that robotic “intelligence” is gradually becoming decoupled from a specific physical body. 2. Long-horizon closed-loop execution. Continuously performing complex tasks for 10 minutes, while being able to pause, switch tasks, and later resume previous ones, is much more meaningful than completing a single 10-second demo. 3. Collaboration and dynamic planning. The two robots appear able to adjust their behavior according to environmental changes and each other’s physical capabilities. These are some of the characteristics that truly general-purpose embodied intelligence will eventually need. That said, I would remain cautious for now. The video demonstrates extremely impressive behavior, but stronger claims such as “self-evolution,” “true understanding of the physical world,” or overturning the Scaling Law with only dozens of hours of training data still require much stronger evidence. Failure recovery and replanning during a task also do not automatically demonstrate that the model is learning by itself. So I wouldn’t call this the “ChatGPT moment” of embodied AI yet. But if the team later discloses the model architecture, training data scale, level of human intervention, and can repeatedly reproduce these capabilities in completely unfamiliar environments, this seemingly rough 10-minute video could become one of the most memorable embodied AI demos of 2026. For now, my biggest question is simple: Who is the mysterious team behind it?

Ice Universe

24,557 Aufrufe • vor 2 Tagen

Yann LeCun (Yann LeCun ) beautifully explains how the architecture and principles used to train LLMs can not be extended to teach AI the real-world intelligence. In 1 line: LLMs excel where intelligence equals sequence prediction over symbols. Real-world intelligence requires learned world models, abstraction, causality, and action planning under uncertainty, which current next-token training does not provide. He says current LLMs learn by predicting the next token. That objective works very well when the task itself can be reduced to manipulating discrete symbols and sequences. Math, physics problem solving on paper, and coding fit this pattern because success largely comes from searching and composing the right sequences of symbols, equations, or program tokens. With enough data and scale, these models get very good at that kind of structured sequence prediction. Real-world intelligence is different. The physical world is continuous, noisy, uncertain, and high dimensional. To act in it, a system needs internal models that capture objects, dynamics, causality, constraints from the body, and the outcomes of actions over time. Humans and animals build abstract representations from rich sensory streams, then make predictions in that abstract space, not at the raw pixel level. That is why a child can learn intuitive physics, plan multi-step actions, and adapt quickly in new situations with little data. His claim about saturation follows from this gap. Scaling token prediction keeps improving symbol manipulation tasks like math and code, but it hits limits on embodied reasoning and common sense because text alone does not provide the right learning signals for world models. Predicting the next word cannot efficiently teach contact forces, affordances, occlusion, friction, or how actions change the state of the environment. For that, he argues we need architectures that learn abstractions from sensory data and predict futures in abstract latent spaces, then use those predictions to plan actions toward goals with built-in guardrails. --- From 'Pioneer Works' YT Channel (link in comment)

Rohan Paul

104,460 Aufrufe • vor 8 Monaten

NVIDIA just unleashed SANA-WM and it’s an absolute MONSTER for the future of open source AI! A blazing-fast 2.6B-parameter open-source world model that doesn’t just generate video… it creates controllable, physics-rich, high-fidelity worlds on demand. Why this is insanely powerful: • One image + text prompt + 6-DoF camera trajectory → generates 720p videos up to 60 seconds long with buttery-smooth, precisely controlled camera movement. You’re not just watching, you’re piloting the simulation. • Runs locally on a single consumer GPU (RTX 5090 level) thanks to heavy distillation + NVFP4 quantization. Full 60-second clip denoised in ~34 seconds. No massive clusters required. • 36× higher throughput than previous open models while rivaling (or beating) closed industrial giants in visual quality and consistency. • Trained lightning-fast: ~213K public videos in just 15 days on 64 H100s. • Built with next-level tech: Hybrid Linear Attention, dual-branch camera control, two-stage pipeline, and rock-solid metric-scale pose understanding. This is a true open world model, the foundation for embodied AI, robotics, autonomous systems, and hyper-realistic simulations that can run anywhere. Project: At our Zero-Human Company, we’re already running SANA-WM live in our core pipelines. It’s supercharging autonomous agent training, generating unlimited synthetic training data, and powering full end-to-end simulation loops, zero humans in the loop. The speed and control let us test thousands of edge-case scenarios overnight, iterate at lightspeed, and push our fully autonomous operations further than ever before. This is the kind of breakthrough that turns science fiction into daily reality. World models just leveled up — hard. The age of personal, local, controllable universes is here.

Brian Roemmele

619,062 Aufrufe • vor 3 Monaten

New framework: Kick down your robot, it will get back up every time 🥋 Chinese startup RoboParty is a Beijing startup founded April 2025 by Huang Yi, originally shipping ROBOTO ORIGIN, the world's first full-stack open-source bipedal humanoid. They released UFO: Unsupervised Reinforcement Learning Framework for Humanoid Control. DEFINITIONS -> what differs is where the learning signal comes from: - SUPERVISED: humans supply the right answers (labels), the model imitates them. - UNSUPERVISED: no answer key, the model finds structure in raw data on its own. - REINFORCEMENT LEARNING: no answer key either, the model tries things and a reward scores each attempt. → UNSUPERVISED RL: trial and error where the agent invents its own rewards, instead of engineers hand-writing one per task. REPRESENTATION LEARNING: compress raw states into a useful internal map. TEMPORAL DISTANCE: distance on that map is "how many steps from A to B." CONTRASTIVE: trained by pulling together what's close in time, pushing apart what isn't. -> CONTRASTIVE TEMPORAL-DISTANCE REPRESENTATION LEARNING: the model builds an internal map of body states where distance means how many steps it takes to get from one to another. It is trained by contrast: states that occur close together in a movement get pulled together in the map, randomly paired states get pushed apart. UFO is an open-source training framework that teaches humanoid robots skills, like getting up, walking, goal-reaching, teleoperation, without reference motions -> no motion-capture or human-video demonstrations to imitate. Its core is TeCH, a contrastive temporal-distance representation-learning algorithm: the robot explores, builds pseudo-goals by temporal rolling, and learns goal-conditioned policies from a single unified progress reward. One framework trains five different robots (Unitree G1/H1, RoboParty RP0/RP1, AgiBot X2) with automatic config conversion in ~2–3 hours per robot! The real novelty here "no demonstrations at all". No data-collection arms race,the dominant humanoid-locomotion recipe is tracking: imitate mocap/retargeted-human reference trajectories. The robot self-generates goals from its own exploration and learns from a progress reward, needing zero reference motion data. Everybody else is fighting over data acquisition, while this team just teleports out of the race entirely (inb4 "competition is for losers 💀 ). This strategy reminds me of the DeepSeek playbook applied to robots: open-source the whole stack to become the global default and commoditize everyone else. RoboParty is giving away hardware and now control software (UFO) to be the Android of humanoids. Yet another reason for the US to ban Chinese open models perhaps 🥶 ? What I also really like about this approach is the cross-embodiment infrastructure, one framework trains Unitree G1/H1, RoboParty RP0/RP1, and AgiBot X2 with automatic configuration conversion. Just like Physical Intelligence, RoboParty seems to place itself as a neutral hardware agnostic middle man. Also woth mentioning: their ability ot perform stable skill injection, e.g. adding a cartwheel without forgetting how to walk. A common failure of RL humanoid policies is that teaching a new agile skill destabilizes the existing ones (catastrophic forgetting). UFO claims you can inject rare motions (cartwheel) without collapsing learned behavior. If it holds, that's a significant incremental/continual skill-learning! But again, I have to underline it: no arXiv, no external validation, no success-rate numbers. -> robotics badely needs an independent unbiased evaluator imho. Still, look at that cool demo: robot is getting kicked and pushed around (serious disturbance) during teleoperation (controlled the person at the back wearing the VR headset), and still managed to always get back up. This is some serious demonstration of stability and robustness!

Léo

35,855 Aufrufe • vor 26 Tagen