Загрузка видео...

Не удалось загрузить видео

На главную

🚧 Progress update: reBot Arm is taking shape.👉 🎯#OpenSource 6+1 DoF force-controlled robot arm ; 🎯Native support for high-end torque #motors - #DAMIAO 4310, RobStride Dynamics 06 QDD; 🎯Deliver smooth joint-level force/position control, open low-level protocols, and #ROS2/#Python compatibility; 🎯Built for bridging #simulation and real-world AI manipulation.

14,918 просмотров • 5 месяцев назад •via X (Twitter)

Комментарии: 0

Нет доступных комментариев

Здесь появятся комментарии из оригинального поста

Похожие видео

🔥 JUST IN: Open-source robotics dataset from 100% real-world scenarios! 🤯 Chinese robotics company AGIBOT just released AGIBOT WORLD 2026, an open-source dataset systematically covering key embodied AI research directions. Built entirely from real-world environments: commercial spaces, and homes. Collected using AGIBOT G2 robots in free-form collection mode, providing structured, accurately annotated, high-quality data. Digital twin technology creates 1:1 scale replicas in simulation matching the real environments. Both real-world and simulation data are open-sourced. The AGIBOT G2 platform collects multiple data types simultaneously: RGB(D) cameras, tactile sensors, force sensors, LiDAR, IMU, and full-body joint states. Whole-body control coordinates arms, waist, and hands for complex tasks. First-person teleoperation lets operators control the robot from its perspective. The tasks covered are fine-grained manipulation, ultra-long-horizon tasks, spatial navigation, dual-arm coordination, and multi-agent/human-robot collaboration. The dataset includes error-recovery trajectories with annotations. Most datasets only show successful demonstrations. AGIBOT includes failures and how the robot recovers, teaching models how to handle mistakes. After collection, data is tested through policy training and real-robot deployment to ensure quality. Then processed through industrial quality control with multiple screening and cleaning rounds. Making it open-source accelerates embodied AI research by giving researchers access to high-quality real-world robot data at scale. 🇨🇳 Learn more here: ~~ ♻️ Join the weekly robotics newsletter, and never miss any news →

Lukas Ziegler

40,583 просмотров • 4 месяцев назад

X-Humanoid just officially dropped Embodied Tien Kung 3.0, A universal platform designed to be way more open and developer-friendly. 🤖 Built on their Wise Kaiwu AI platform, this next-gen humanoid is all about slashing development costs. It’s a fully interoperable ecosystem that supports everything from tactile interaction to high-dynamic motion control at a full humanoid scale. ➤ Radical Openness: X-Humanoid is open-sourcing the full stack—robot body, motion control, VLM/VLA models, and the RoboMIND dataset. It fully supports ROS2, MQTT, and TCP/IP, so developers can customize use cases without re-engineering the basics. ➤ High-Performance Hardware: With high-torque integrated joints, Tien Kung 3.0 can clear 1-meter (3.3ft) obstacles and handle dexterous moves like kneeling and bending. It hits millimeter-level precision, making it a solid fit for industrial-grade tasks. ➤ True Autonomy: The bot runs a continuous perception-decision-execution loop. It uses world models to break down complex language commands and VLA models for real-time obstacle avoidance and navigation. ➤ Scalable Collaboration: The platform moves beyond single-unit tasks to support multi-robot collaboration with autonomous scheduling. It’s built to move embodied AI from the lab straight into real-world commercial and industrial environments. Source: X-Humanoid #Humanoid #OpenSource #Robotics #EmbodiedAI #PhysicalAI #Automation #XHumanoid #TienKung #WiseKaiwu

RoboHub🤖

49,440 просмотров • 6 месяцев назад

Today, we're joined by Nikita Rudin, co-founder and CEO of Flexion to discuss the gap between current robotic capabilities and what’s required to deploy fully autonomous robots in the real world. Nikita explains how reinforcement learning and simulation have driven rapid progress in robot locomotion—and why locomotion is still far from “solved.” We dig into the sim2real gap, and how adding visual inputs introduces noise and significantly complicates sim-to-real transfer. We also explore the debate between end-to-end models and modular approaches, and why separating locomotion, planning, and semantics remains a pragmatic approach today. Nikita also introduces the concept of "real-to-sim", which uses real-world data to refine simulation parameters for higher fidelity training, discusses how reinforcement learning, imitation learning, and teleoperation data are combined to train robust policies for both quadruped and humanoid robots, and introduces Flexion's hierarchical approach that utilizes pre-trained Vision-Language Models (VLMs) for high-level task orchestration with Vision-Language-Action (VLA) models and low-level whole-body trackers. Finally, Nikita shares the behind-the-scenes in humanoid robot demos, his take on reinforcement learning in simulation versus the real world, the nuances of reward tuning, and offers practical advice for researchers and practitioners looking to get started in robotics today. 🗒️ For the full list of resources for this episode, visit the show notes page: 📖 CHAPTERS =============================== 00:00 - Introduction 04:07 - Is robot locomotion solved? 06:04 - Sim-to-real gap 08:58 - Adding semantics to policies 09:42 - Modular vs end-to-end architectures 10:29 - Planner model 12:21 - Adapting RL techniques from quadrupeds to humanoids 15:39 - Behind robot demos 18:09 - Humanoid robots in home environments 22:03 - Training approach 23:56 - VLA models 27:59 - Closing the sim-to-real gap 32:55 - Task orchestration using VLMs 36:38 - Tool use 38:10 - Model hierarchy 43:37 - Simulator versus simulation environment 44:57 - Combining imitation learning and reinforcement learning 46:42 - RL in real world versus RL in simulation 52:58 - Reward tuning and value functions in robotics 56:38 - Predictions 1:00:10 - Humanoids, quadropeds, and wheeled platforms 1:02:45 - Advice, recommended robot kits, and community pla

The TWIML AI Podcast

22,592 просмотров • 7 месяцев назад

NVIDIA just unleashed SANA-WM and it’s an absolute MONSTER for the future of open source AI! A blazing-fast 2.6B-parameter open-source world model that doesn’t just generate video… it creates controllable, physics-rich, high-fidelity worlds on demand. Why this is insanely powerful: • One image + text prompt + 6-DoF camera trajectory → generates 720p videos up to 60 seconds long with buttery-smooth, precisely controlled camera movement. You’re not just watching, you’re piloting the simulation. • Runs locally on a single consumer GPU (RTX 5090 level) thanks to heavy distillation + NVFP4 quantization. Full 60-second clip denoised in ~34 seconds. No massive clusters required. • 36× higher throughput than previous open models while rivaling (or beating) closed industrial giants in visual quality and consistency. • Trained lightning-fast: ~213K public videos in just 15 days on 64 H100s. • Built with next-level tech: Hybrid Linear Attention, dual-branch camera control, two-stage pipeline, and rock-solid metric-scale pose understanding. This is a true open world model, the foundation for embodied AI, robotics, autonomous systems, and hyper-realistic simulations that can run anywhere. Project: At our Zero-Human Company, we’re already running SANA-WM live in our core pipelines. It’s supercharging autonomous agent training, generating unlimited synthetic training data, and powering full end-to-end simulation loops, zero humans in the loop. The speed and control let us test thousands of edge-case scenarios overnight, iterate at lightspeed, and push our fully autonomous operations further than ever before. This is the kind of breakthrough that turns science fiction into daily reality. World models just leveled up — hard. The age of personal, local, controllable universes is here.

Brian Roemmele

618,611 просмотров • 2 месяцев назад

AgiBot’s new generation of industrial-grade interactive embodied robot, AgiBot G2, has officially launched! The G2 has already secured orders worth hundreds of millions of RMB, including two separate contracts each exceeding 100 million RMB, and has begun its first commercial deliveries. The AgiBot G2 is built to industrial standards, featuring high-performance joints, precision torque sensors, and an advanced spatial perception system. It supports rapid learning and deployment, offers strong multimodal voice interaction, and is designed for general use in industrial, logistics, and guidance scenarios. Inheriting the successful "Collect-Train-Deploy" model of its predecessor, the G1, the G2 brings significant upgrades, including a high-performance AI computing platform and actuators that enable omnidirectional obstacle avoidance and high-precision force-control tasks. Its 3-DOF waist allows for human-like bending and lateral body movement. A key feature is the G2's globally first-of-its-kind cross-shaped wrist force-control arm, which uses precision joint torque sensors and joint impedance control to delicately perceive external forces and respond smoothly. For continuous operation, the G2 supports autonomous charging and features a dual-battery hot-swapping system, meeting the 24-hour cycle demands of factory production lines. During the launch event, AgiBot demonstrated the G2’s ultra-low latency remote operation (teleoperation) capabilities. Operators successfully demonstrated precision shots (like hitting a floating balloon in Shanghai while operating from Beijing), showcasing the robot's high accuracy and low latency in both line-of-sight and beyond-line-of-sight scenarios. The G2 is already being deployed across four key real-world scenarios: In automotive parts production, it assists humans with tasks like safety belt lock core pressing and material handling. In precision operations, it used reinforcement learning to master delicate tasks like inserting memory sticks in just one hour. In logistics, the G2, enhanced by AgiBot's OmniHand dexterous hand, efficiently handles various package types for sorting and loading. Its strong mobility allows it to adapt to over 95% of factory floors. AgiBot is also commencing the first batch of commercial deliveries under an over 100 million RMB procurement contract with Joyson Electronic, formally landing the G2 in the automotive parts manufacturing sector.

RoboHub🤖

33,831 просмотров • 10 месяцев назад

China unveils humanoid robot worker with brain that runs 275 trillion ops/sec | Jijo Malayil, Interesting Engineering In tests, SUYUAN used vision and joint control to sort and move crates of various sizes, greatly improving warehouse productivity. Chinese manufacturing firm Shanghai Electric has unveiled its first self-developed industrial humanoid robot, “SUYUAN,” marking a major milestone in its robotics journey. Debuting at the World Artificial Intelligence Conference (WAIC 2025) on July 26 in Shanghai, SUYUAN boasts 38 degrees of freedom and 275 TOPS of on-device computing power, enabling precise operations and fluid movements. According to the firm, designed for diverse industrial use, the robot showcases Shanghai Electric’s end-to-end capabilities—from core tech to integrated solutions—and reinforces its commitment to next-gen industrial automation through a full industry chain strategy. At WAIC 2025, Shanghai Electric also unveiled a new joint venture with Johnson Electric for next-gen humanoid robotics and showcased its “LINGKE” dual-arm robot. Recently, Hangzhou-based Unitree Robotics launched the R1 humanoid with 26 joints for $5,900, showcasing athletic feats like cartwheels, running, and quick recovery. Smart factory assistant Shanghai Electric claims SUYUAN, equipped with 38 degrees of freedom (DoF) and a powerful 275 TOPS on-device computing processor, delivers fluid, human-like movements and high-precision operations across various industrial scenarios. Its advanced articulation and real-time processing capabilities make it highly adaptable, enabling smooth execution of complex tasks in dynamic work environments. SUYUAN, who weighs 110 pounds (50 kilograms) and is 5 feet 6 inches (167 cm) tall, was designed to have human-like proportions. Its 38-DoF articulation offers dexterity, allowing for both wide-range motion and sensitive manipulation. With a single arm, the robot can lift objects up to 4.4 pounds (2 kilograms) in weight and carry a total payload of up to 22 pounds (10 kilograms). With a walking pace of 3.1 miles per hour (5 km/h), SUYUAN is ideal for environments including assembly lines, warehousing, and logistics, according to a statement. To navigate complex industrial settings, SUYUAN combines LiDAR and binocular vision for self-guided mobility. Its 275-TOPS AI processor enables rapid data analysis and integration with large language models, allowing it to understand tasks in natural language and handle objects adaptively, reports Fox 44 News. In pilot demonstrations, the robot successfully identified, picked, and relocated crates of varying sizes using advanced computer vision and coordinated joint control—delivering measurable gains in warehouse efficiency. The company claims that SUYUAN’s launch represents a major turning point in Shanghai Electric’s foray into humanoid robotics and strengthens its vertically integrated approach to industrial automation solutions. Intelligent task handling Shanghai Electric also demonstrated its most recent developments in intelligent manufacturing at WAIC 2025, introducing a new joint venture with Johnson Electric centered on next-generation humanoid robotics and showcasing the “LINGKE” dual-arm robot. With its high-precision operations, adaptive teamwork, and closed-loop data capabilities, the LINGKE robot demonstrated live talents in handling complicated production jobs. LINGKE is made to do more than just replace human labor; it uses compliant force control and bimanual coordination to relieve workers of high-intensity, repetitive jobs. According to the company, the robot enhances operational efficiency by up to five times. Its core strength lies in a Data-Model-Deployment closed-loop system that starts with operational data, followed by data cleansing, model training, live deployment, and feedback-driven optimization—enabling autonomous learning and workflow improvement. Also at the event, Shanghai Electric and Johnson Electric introduced advanced hardware modules for humanoid robots, including rotary joints, linear joints, and dexterous finger joints. These components are designed to support smooth, precise, and quiet motion performance across robotics systems, reports Stock Titan. The joint venture announced two strategic agreements: a first-unit supply deal with the National and Local Co-Built Humanoid Robotics Innovation Center (Qinglong Project) and a cooperation memorandum with Fourier Robotics. Read more:

Owen Gregorian

51,638 просмотров • 1 год назад

China unveils humanoid robot with lifelike skin and blinking eyes built for daily life | Prabhat Ranjan Mishra, Interesting Engineering Large Language Models (LLMs) and Vision-Language Models (VLMs) help process and interpret complex data from human interactions. A Shanghai-based company has developed humanoid robots that appear as real as humans. The advanced bionic humanoid robot is integrated with self-supervised AI algorithms. Named Elf V1, the robot can perceive the world, communicate, learn, and interact intelligently with its surroundings. Developed by AheadForm Technology, the robot offers up to 30 degrees of freedom, powered by a precise control system and an advanced AI learning algorithm. Robot offers expressive facial features The robot offers expressive facial features, moving eyes, and synchronized speech. It can also convey emotions and understand human non-verbal cues, making interactions more natural and engaging. The robot has highly interactive capabilities and lifelike appearances. AheadForm expects that its robots could soon seamlessly integrate into daily life, providing assistance, companionship, and support across various industries. “We believe that by developing realistic and expressive robot heads, we can bridge the gap between humans and machines, fostering a new era of interactive and intelligent robotics,” said the company in a statement. Reports revealed that to avoid the “uncanny valley” effect and be able to interact with us, they are given lifelike skin and capabilities to read our emotions and respond appropriately using dynamic expression simulation and emotion generation tech. Bionic skin and high-precision control system The Elf V1 series of humanoids features 30 facial muscles animated by brushless micro-motors and managed by a high-precision control system. Paired with an ability to detect their users’ emotions with low latency and bionic skin, their facial expressions are nearly identical to those of humans, reported CGTN. The company claims it’s pioneering the development of realistic humanoid robots designed to revolutionize human-robot interaction. It’s enhancing sophisticated humanoid robot heads that can express emotions, perceive their environment, and interact seamlessly with humans. By combining cutting-edge AI and advanced robotics, AheadForm aims to bring life to machines and transform how humans engage with technology. AI models boost robots’ responsiveness Seamless integration of Large Language Models (LLMs) and Vision-Language Models (VLMs) into the humanoid robots can help them process and interpret complex data from human interactions, enabling the robot to learn and adapt in real-time, achieving human-level understanding and responsiveness. AheadForm uses Brushless Motors that deliver ultra-quiet operation and high responsiveness, specifically designed for precision facial movements in humanoid robots. With its compact size, lightweight design, and energy efficiency, this motor is the ideal choice for next-generation robots that require precise, subtle facial control to create a truly human-like experience. Previously, the company unveiled the Lan Series that features realistic humanoid robots with soft skin and 10 degrees of freedom, offering a lifelike appearance and intuitive movements. This series is designed for cost-efficiency, for applications prioritizing mobility and manipulation.

Owen Gregorian

179,005 просмотров • 9 месяцев назад

Excited to announce GR00T N1, the world’s first open foundation model for humanoid robots! We are on a mission to democratize Physical AI. The power of general robot brain, in the palm of your hand - with only 2B parameters, N1 learns from the most diverse physical action dataset ever compiled and punches above its weight: - Real humanoid teleoperation data. - Large-scale simulation data: we are open-sourcing 300K+ trajectories! - Neural trajectories: we apply SOTA video generation models to “hallucinate” new synthetic data that features accurate physics in pixels. Using Jensen’s words, “systematically infinite data”! - Latent actions: we develop novel algorithms to extract action tokens from in-the-wild human videos and neural generated videos. GR00T N1 is a single end-to-end neural net, from photons to actions: - Vision-Language Model (System 2) that interprets the physical world through vision and language instructions, enabling robots to reason about their environment and instructions, and plan the right actions. - Diffusion Transformer (System 1) that “renders” smooth and precise motor actions at 120 Hz, executing the latent plan made by System 2. We deploy N1 on GR1 robot, 1X Neo robot, and a large collection of simulation benchmarks. N1 achieves up to +30% boost in diverse manipulation tasks for household and industrial settings. While humanoid robots are the main focus of N1, our model also supports cross-embodiment. We finetune it to work on the $110 HuggingFace LeRobot SO100 robot arm! Open robot brain runs on open hardware. Sounds just right. Let’s solve robotics, together, one token at a time. Links to our Whitepaper, Github repo, HuggingFace model, and open dataset page in the thread: 🧵

Jim Fan

466,442 просмотров • 1 год назад

Most robotics AI models suffer from the "stop-and-think" problem. They take a static picture, pause to reason, execute an action, and repeat. In the real world, that latency causes spills, collisions, and failed tasks. Google DeepMind just launched Gemini Robotics ER 2: an embodied reasoning model that thinks and acts at the speed of the physical world. Here's why this is a step-change for physical AI engineering: Traditional robotics models rely on static snapshots. But knowing *when* a task is done, such as when to stop pouring coffee into a cup or when a trash bag is securely tied, requires continuous temporal awareness. Gemini Robotics ER 2 integrates directly with the bidirectional streaming Gemini Live API to reason about what comes next while simultaneously executing motor actions. What makes Gemini Robotics ER 2 different: 🎯 91.3% accuracy on live video moment-finding (0.96s mean absolute distance) at 4x the execution speed of frontier models 📈 Continuous progress tracking across 5 completion stages (57.4% accuracy) to self-correct mid-task without restarting 🛠️ Native agentic tool orchestration that commands lower-level VLA models, navigation APIs, and Google Search 🤝 Multi-robot collaboration allowing physically diverse machines (like Apptronik's Apollo 2 humanoid and Franka's FR3 Duo arm) to hand off tasks in shared spaces 🛡️ Built-in physical safety that autonomously halts robots when humans enter a workspace and resumes once clear

Karl Weinmeister

28,107 просмотров • 13 дней назад

Most video-action robot models are a content-creation video generator with an action module attached. LingBot-VA 2.0 from Robbyant, a video-action foundation model, throws that starting point out and trains the whole stack natively for control. And it runs closed-loop at a peak 225 Hz. It's so important because A robot cannot move responsively when its controller pauses to imagine the next few frames. LingBot-VA 2.0 predicts during execution, then corrects using each real observation. And it carries only about 13B video parameters while activating roughly 1.9B per token. Bigger robot models usually mean slower reactions, creating a direct conflict between intelligence and control. LingBot-VA 2.0 is trained from scratch for robot control rather than adapted from a video generator built for content creation. Robbyant, an embodied AI company under Ant Group, built it to learn how scenes change under actions, predict what should happen next, and turn those predictions into real-time robot movements. Most video-action systems inherit a tokenizer and video backbone trained mainly to reproduce visual appearance. LingBot-VA 2.0 rebuilds both parts around physical control. Its semantic visual-action tokenizer maps observations toward features from a frozen vision foundation model and learns compact latent actions from frame-to-frame changes using self-supervised inverse and forward dynamics. Unlabeled web video can therefore carry action-relevant training signals without robot action labels. The policy is causal from the start, so every prediction can use only past observations. Its sparse Mixture-of-Experts video backbone has about 13B total parameters, while about 1.9B are active per token, keeping the compute lower during each step. A high-level vision-language planner breaks long tasks into smaller instructions, while the low-level video-action policy handles continuous movement. Foresight Reasoning predicts future visual states while the robot is already acting, then replaces imagined states with every new real observation. Combined with few-step distillation and systems acceleration, the paper reports a peak asynchronous execution frequency of 225 Hz. The model adapts from 10–15 demonstrations, transfers across robot embodiments, and handles some new tasks zero-shot. In the paper’s own evaluations, it reaches 93.6 average on RoboTwin 2.0 and reports stronger real-world results than LingBot-VA and π0.5 across the tested tasks. 🧵 1.

Rohan Paul

11,253 просмотров • 1 месяц назад

🚀 Introducing PantheonOS ( A Fully Open-Source Agent OS for Science PantheonOS began as a research project in my Stanford lab and has since evolved into a vision to redefine data science in the era of AI—starting with computational biology, especially single-cell and spatial genomics. PantheonOS is a general agent platform built from the ground up. It is arguably the first distributed agent framework designed for scientific data analysis. 🔑 Key Features 1. Multi-Agent Collaboration – Built-in paradigms for distributed, cross-machine cooperation among agents and toolsets. 2. Native Toolset Support – Python, R, Julia, LaTeX, and more—designed for real scientific workflows. 3. Modular & Extensible – Developer-friendly design with shallow wrappers, plus LLM-driven toolset generation. 4. Evolvable Agents – Capable of evolving large-scale code projects to achieve superhuman performance (e.g., evolving upon the original Harmony [I Korsunsky, 2019, Nature Biotechnology] and Scanorama [BL Hie, 2019, Nature Biotechnology] implementations), and even evolving the system itself to adapt to new fields. 🎉 Stepwise Release Strategy We’re releasing PantheonOS in stages: Pantheon-CLI (today!), followed by Pantheon-Lab, Pantheon-Notebook, Pantheon-Slack, and more. 🌟 Pantheon-CLI Highlights - We're not just building another CLI tool. We're defining how scientists will interact with data in the AI era. - Open, Powerful, Python-First – The first fully open-source, endlessly extendable scientific “vibe analysis” framework. - Mixed Programming Magic – Combine Python, natural language, R, or Julia—seamlessly in the same environment. - PhD-Level Assistant – A command-line agent for complex real-world genomics and beyond, handling workflows at the PhD level. - Privacy by Design – Run entirely offline with local LLMs—your data never leaves your computer. ✅ Proven Applications (10 Demonstrations) Computational biology: 1. ATAC-seq: From raw reads to peak matrix 2. RNA-seq: From raw reads to expression matrix 3. Complex single-cell workflows (PhD-level) 4. Hybrid natural language + R for Seurat annotation 5. Learning from web tutorials + invoking single-cell foundation models 6. Cell segmentation on 10x Genomics HD Visium data And beyond: 7. Mixed Python & R programming examples 8. Molecular docking & structural analysis 9. Exploratory factor analysis for behavioral survey data 10. Customer segmentation & finance analytics 🌐 Learn More & Get Started Website: Pantheon-CLI Documentation: GitHub Repo: 💬 Join our community: PantheonOS Slack: PantheonOS Discord:

evo-devo

17,384 просмотров • 11 месяцев назад

My conversation with Sergey Levine (Sergey Levine). Sergey is the co-founder of Physical Intelligence -- a company building foundation models that can control any robot to do any task in any environment. The company's thesis is that generality is more scalable than specialization, meaning that a model trained across many different robots and tasks will ultimately outperform any system built to do one thing well (eg, just wash dishes). Sergey is a researcher by background, but I think you will appreciate how practical and commercially grounded this conversation is. We discuss: - Why changing a diaper will be the last task a robot masters - The simulation v. real-world data debate - How multimodal LLMs give robots common sense - Moravec's Paradox + Robot Olympics - Why robots can do long-horizon tasks now - A realistic timeline for robots in our homes I should note that I am an investor in Physical Intelligence -- I made the investment because I believe it is one of the most important companies tackling the problem of robotics. Enjoy! Timestamps: 0:00 Intro 2:39 Defining Physical Intelligence 5:19 The Challenge of Building General Models 6:34 The Stakes and Future of General Purpose Robotics 8:15 Pros and Cons of Humanoid Robots 10:12 Historical Milestones in Robotics Research 15:31 Combining Generative AI and Deep RL 21:24 Moravec's Paradox 25:33 Kitchen Robots 29:30 Simulation vs. Real-World Data 30:48 The Robot Olympics 36:31 The Physiological Reality of Embodiment 38:56 Controversies in the Robotics Community 44:18 What Makes a Great Researcher 48:27 How Businesses Should Prepare for Robotics 54:09 Tracking Progress Through Research Papers 57:02 The Next Step: Mid-Level Reasoning 1:02:00 The Kindest Thing

Patrick OShaughnessy

133,833 просмотров • 4 месяцев назад

NEW RESEARCH: You can now create a new robot optimized for any given task! I love this new project by Huy Ha, Shuran Song, and others. Called "Transformer Transformer: A Unified Model for Motion-Conditioned Robot Co-design", it generates a robot's physical design and its controller together from a task spec. DEFINITIONS: - Reward function: A scoring rule that assigns a number to how well a behavior achieves the task. Here, it is the objective the generated design is pushed to maximize (e.g., track the target motion with low error). - Tokenizing: dividing continuous or structured data (a robot's links, joints, motor specs, states, actions) into a discrete vocabulary of symbols a transformer can process, the same step that turned pixels and audio into "language" for these models. - Diffusion transformer (DiT): A transformer trained to turn random noise into structured output through iterative denoising. Here, it generates robot bodies and trajectories instead of images. - MuJoCo: The standard fast physics simulator for robotics research (DeepMind-maintained). The Menagerie is its curated zoo of ready-to-use robot models. - CMA-ES: Covariance Matrix Adaptation Evolution Strategy, the workhorse black-box optimizer: it evolves a population of candidate designs, keeps the best, and needs thousands of simulator rollouts. - Bimanual multi-trajectory optimization: Finding one design/controller that performs well across several target motions for a two-armed robot at once, harder than optimizing for a single arm and a single motion. - BERT/MAE masked-modeling trick: Train one model to fill in whatever parts of the input you hide (words for BERT, image patches for MAE); at inference, choosing what to mask chooses the task, so masking the body makes it a designer and masking the actions makes it a controller. In practice, you give it a target end-effector motion and a reward function, and it outputs a complete embodiment (link, joint, motor, and inertial property), as well as a controller to drive it. It works by tokenizing both the body (links/joints/motors) and the dynamics (states/actions) into a compact scheme called RoboTokens, training a diffusion transformer (DiT) over them. The same model predicts dynamics using those predictions ("Dynamics Self-Guidance") to push generated designs toward higher reward at inference time. Masking different token types (using the BERT/MAE masked-modeling trick) lets the one model do three jobs: generate an embodiment, control an arbitrary embodiment, or design one conditioned on a motion. It is trained on 11 robots from the MuJoCo Menagerie (0.65 kg hand to 67.5 kg quadruped, 6–35 joints), and validated in sim and on a physical ALOHA doing cloth flinging. I like the fact that this approach inverts the entire recent robotics ideas: designing a policy for a fixed robot -> designing the robot for a fixed task. Every other approach assumes the body is given and learns a controller. Transformer Transformer takes the task (target motion + reward), then generates the body and controller jointly. In practice, it is a ~180× speedup over the standard optimizer at equal-or-better quality. It reaches "CMA-ES-level quality in seconds" and finishes bimanual multi-trajectory optimization in that is worth underlining nowadays! Also worth mentioning: this is the lab behind UMI and Handroid, that I mentioned here previously! The team seems extremely creative, i love these out-of-the-box approaches. Enjoy watching the demo of robot optimization in 3D, data acquisition, then real-life testing:

Léo

25,680 просмотров • 7 дней назад

China’s pretty humanoid robot stuns by opening a car door in a ‘world’s first’ | Jijo Malayil, Interesting Engineering Mornine used onboard sensors and full-body control to locate the handle, adjust posture, and open a car door—no human input needed. AiMOGA Robotics has claimed to have reached a significant milestone in embodied AI with its humanoid robot, Mornine, autonomously opening a car door inside a functioning Chery dealership in China. Relying solely on onboard sensors, full-body motion control, and end-to-end reinforcement learning, Mornine performed the task without any human input. Unlike scripted or teleoperated robots, Mornie identified the door handle, adjusted its posture, and used coordinated force across its limbs and torso to complete the action—demonstrating advanced autonomy in a real-world setting. “The deployment marks one of the first instances of a service robot executing such a high-friction, physical interaction in a live commercial setting,” said the firm in a statement. In April, at the Shanghai Auto Show, automotive brands Omoda and Jaecoo, subsidiaries of Chery Automobile, introduced Mornine, designed for use in car dealerships. From sim to service Opening a car door may seem like a simple task, but AiMOGA Robotics views it as a pivotal moment in robotics—signaling a shift from simulation to real-world service, and from basic command execution to autonomous capability. Using only onboard sensors and full-body motion control, Mornine identified the door handle, adjusted her posture, and applied coordinated force across her limbs to open the door—entirely without human intervention. Mornine’s advanced sensor suite includes 3D LiDAR, depth and wide-angle cameras, and a visual-language model (VLM), enabling real-time perception of door position and opening status. Uniquely, Mornine wasn’t explicitly programmed to recognize door handles. Instead, she learned through reinforcement learning, undergoing millions of simulated cycles to focus on the right region and perform the task independently. “We never explicitly told the robot what a door handle is. It learned to focus on that region by itself,” said the engineering team at AiMOGA Robotics in a statement. The learned model was transferred to the real world using Sim2Real methods. Mornine continuously gathers live sensor data during operation, which feeds into a cloud-based training loop, allowing her to improve through continuous learning in real-world settings, reports Robotics Tomorrow. Now active in multiple Chery 4S dealerships in China, Mornine not only opens car doors but also assists with customer greetings, vehicle introductions, and item delivery—marking a step forward in humanoid robotics for commercial retail environments. AI meets retail Originally introduced as the AiMOGA Robot, Mornine was developed to support dealership sales by performing tasks such as explaining vehicle specifications, leading showroom tours, serving refreshments, and engaging with customers in multiple languages. First conceived by Chery as a virtual character to appeal to Generation Z using metaverse and virtual human technologies, Mornine gradually evolved into a real-world interactive humanoid. After multiple iterations of character and model design, Mornine debuted as a digital persona in animations, livestreams, and promotional content, gaining brand recognition. Chery later expanded the concept beyond the virtual space, resulting in the creation of the AiMOGA humanoid robot. Leveraging Chery’s expertise in autonomous driving, environmental sensing, and control systems, AiMOGA features full-stack capabilities in perception, cognition, decision-making, and execution. It uses multimodal sensing—combining speech, vision, and environmental data—to interpret user gestures, commands, and showroom dynamics. A bionic motion system and automotive-grade hardware enable dexterous movement and upright mobility, while multi-robot collaboration allows for coordinated tasks like guided tours. At the decision-making layer, Deepseek’s large language models enable natural language understanding and personalized interaction. In April 2025, Mornine officially began commercial service as an “Intelligent Sales Consultant” at the OMODA C5 JOYSTAR 4S dealership in Kuala Lumpur, Malaysia—marking her full transition from a virtual concept to a real-world humanoid sales assistant.

Owen Gregorian

67,975 просмотров • 1 год назад