Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Introducing TRLC-DK1-X: An accessible humanoid data collection platform. It features two 6+1 DOF leader-follower arm pairs and three ultra-wide FOV cameras, integrated for learning static, bimanual manipulation tasks. Order now for $4,999.

44,286 Aufrufe • vor 10 Monaten •via X (Twitter)

0 Kommentare

Keine Kommentare verfügbar

Kommentare vom Original-Post werden hier angezeigt

Ähnliche Videos

🔥 JUST IN: Open-source robotics dataset from 100% real-world scenarios! 🤯 Chinese robotics company AGIBOT just released AGIBOT WORLD 2026, an open-source dataset systematically covering key embodied AI research directions. Built entirely from real-world environments: commercial spaces, and homes. Collected using AGIBOT G2 robots in free-form collection mode, providing structured, accurately annotated, high-quality data. Digital twin technology creates 1:1 scale replicas in simulation matching the real environments. Both real-world and simulation data are open-sourced. The AGIBOT G2 platform collects multiple data types simultaneously: RGB(D) cameras, tactile sensors, force sensors, LiDAR, IMU, and full-body joint states. Whole-body control coordinates arms, waist, and hands for complex tasks. First-person teleoperation lets operators control the robot from its perspective. The tasks covered are fine-grained manipulation, ultra-long-horizon tasks, spatial navigation, dual-arm coordination, and multi-agent/human-robot collaboration. The dataset includes error-recovery trajectories with annotations. Most datasets only show successful demonstrations. AGIBOT includes failures and how the robot recovers, teaching models how to handle mistakes. After collection, data is tested through policy training and real-robot deployment to ensure quality. Then processed through industrial quality control with multiple screening and cleaning rounds. Making it open-source accelerates embodied AI research by giving researchers access to high-quality real-world robot data at scale. 🇨🇳 Learn more here: ~~ ♻️ Join the weekly robotics newsletter, and never miss any news →

Lukas Ziegler

40,583 Aufrufe • vor 5 Monaten

🦔Workers in India are wearing head-mounted cameras for 12 cents an hour to collect training data for humanoid robots. The footage of them doing everyday tasks like cooking, cleaning, sorting, and walking through public spaces gets sold to robotics companies building the models meant to replace those same kinds of jobs in higher-wage countries. The arrangement has been running for roughly two years. Workers do not own the data, do not get residuals, and in many cases are not told what their footage is being used to train. My Take The workers wearing the cameras live in a country where robotics automation will hit decades later, so they are training their own future replacements at a delay that hides the consequence from them personally. The companies buying the data are mostly US and Chinese, building humanoid robots aimed at warehouses, retail, and service jobs in countries paying $15 to $25 an hour rather than 12 cents. Robotics companies need motion data that mimics how humans actually move through real environments, and synthetic data has not been good enough yet. Paying 12 cents an hour in Bengaluru is cheaper than running motion capture studios in Boston, and it works at scale because the worker absorbs the cost of the camera, the discomfort of wearing it, and the long-term loss of any rights to their own movement data. The robotics labor market that eventually emerges from this footage will displace far more wages than the data collection cost to gather. That is the trade investors funding humanoid robotics startups are betting will pay off, and the workers in the videos are the ones paying the tab up front. Hedgie🤗

Hedgie

162,060 Aufrufe • vor 3 Monaten

🚨 BREAKING: Microsoft's first robotics foundation model! 🤯 Microsoft just announced Rho-alpha (ρα), their first robotics model derived from the Phi series of vision-language models. Rho-alpha translates natural language commands into control signals for robotic systems performing bimanual manipulation tasks. Commands like "push the green button with the right gripper," "pull out the red wire," "flip the top switch on," or "turn the knob to position 5" get executed directly by dual-arm robots. What makes this different from standard vision-language-action (VLA) models is the additional modalities. Rho-alpha is a VLA+ model that adds tactile sensing to the perceptual mix, with plans to incorporate force feedback. On the learning side, the model is designed to continually improve during deployment by learning from human feedback. The training approach combines trajectories from physical demonstrations and simulated tasks with web-scale visual question answering data. Since teleoperation data is scarce and expensive, Microsoft is using NVIDIA Isaac Sim on Azure to generate physically accurate synthetic datasets via reinforcement learning. These simulated trajectories get combined with commercial and open physical demonstration datasets. The model is currently under evaluation on dual-arm setups and humanoid robots. Microsoft is opening an Early Access Program for organizations interested in evaluating Rho-alpha. Robots that can adapt to dynamic situations and human preferences are more useful in real environments and more trusted by the people operating them. Read more here: ~~ ♻️ Join the weekly robotics newsletter, and never miss any news →

Lukas Ziegler

60,985 Aufrufe • vor 7 Monaten

Humanoid Marathon Champion "TienKung" Enters the Factory At IROS 2025, the Beijing Humanoid Robot Innovation Center (BHRIC) and UBTECH officially launched the Software Development Kit (SDK) for their general-purpose embodied intelligence platform, "HuiSiKaiWu." This move signals a significant step towards an open-source embodied AI ecosystem. The HuiSiKaiWu platform is designed for ease of use and low-threshold deployment, enabling multi-agent collaboration (one brain, multiple functions/robots). The core "brain" uses a dual-model architecture: the Pelican VLM and the WoW World Model, driving autonomous learning and decision-making. The "cerebellum" includes the cross-body XR-1 VLA Model, which has shown strong performance in rapid, few-shot skill transfer across tasks. BHRIC simultaneously announced its first industrial application: since September 2025, the "TienKung 2.0" and "Tianyi 2.0" humanoids have been deployed at the Foton Cummins engine factory. They autonomously handle material bin fetching and transportation on the "unmanned production line," adapting to various goods and shelf heights. Foton Cummins highlighted the robots' high stability and generalization potential in complex industrial processes. Beyond industrial work, BHRIC showcased other applications: the high-performance "TienKung Ultra" is being used at the Li-Ning Sports Science Lab for running shoe tests, precisely simulating human gaits. The platform is expanding applications across manufacturing, logistics, and sports science. The HuiSiKaiWu SDK is now available for collaborative trials by research institutions and development teams, offering a full toolchain for skill calling and scene deployment. BHRIC plans to gradually open-source its core algorithms, including VLM and VLA models, and hosted the "X-Humanoid Young Talent Meetup" and the "2026 Re-Action Humanoid Robot Challenge" to build a comprehensive developer ecosystem.

RoboHub🤖

39,337 Aufrufe • vor 10 Monaten

Force-sensing fingers! 🧤 Stanford researchers just released UMI-FT, a handheld data collection platform that puts compact six-axis force/torque sensors on each finger, enabling finger-level wrench measurements alongside RGB, depth, and pose data. Many manipulation tasks require careful force modulation: too little force and the task fails, too much and you cause damage. But commercial force/torque sensors are expensive, bulky, and fragile, which has limited large-scale force-aware policy learning. UMI-FT changes the economics. The platform uses an iPhone for RGB vision, ultrawide RGB, depth, and pose via ARKit, with each finger sensorized using a CoinFT sensor to capture per-finger wrench information during manipulation. This multimodal data trains an adaptive compliance policy that predicts position targets, grasp force, and stiffness for execution on standard compliance controllers. The learned policy runs slowest and generates reference targets, while model-based compliance and force controllers provide delicate 6D compliance control and real-time force modulation. They tested on three contact-rich, force-sensitive tasks: whiteboard wiping (locate eraser, grasp, wipe until clean), skewering zucchini (grasp slice firmly, push onto stick until punctured), and lightbulb insertion (grasp bulb, align bayonet pin with socket slit, insert while overcoming spring force, rotate to light up). The results are clear. Policies without compliance struggle to modulate contact force and trigger safety faults from excessive force. Policies without force sensing fail to grasp unseen objects or resist reaction forces, causing slippage. Here's the project page: ~~ ♻️ Join the weekly robotics newsletter, and never miss any news →

Lukas Ziegler

12,822 Aufrufe • vor 7 Monaten

This is how ALOHA's "teleoperation" system works - a fancy word for "remote control". Training robots will be more and more like playing games in the physical world. A human operates a "joystick++" to perform tasks and collect data, or intervene if there's any safety concern. There's actually a learning curve to master the controller, much like practicing gaming skills. Teleoperation can be done in many different ways. ALOHA is an impressive custom-built system with very low cost. Here're a few alternatives: (1) Motion Capture (MoCap): apply the MoCap systems used for Hollywood movies to capture the fine-grained motions of hand joints. There would be no "embodiment gap" if the robot hand has 5 fingers. For instance, a demonstrator can wear a CyberGlove ( and manipulate the objects. CyberGlove will capture the motion signals & haptic feedback in real-time, which can be re-targeted onto the humanoid. (2) Wearing gloves & markers can be clumsy. An alternative way to do MoCap is through computer vision. DexPilot from NVIDIA enables marker-less and glove-free data collection. The human operator simply uses their bare hands to perform the tasks. 4 Intel RealSense depth cameras and 2 NVIDIA Titan XP GPUs (yeah, 2019 work) translate the pixels to precise motion signals for robot learning. (3) VR Headset: turn the training room into a VR game and "role play" the robot. This has the advantage of scalable remote data collection - annotators from around the world can contribute without coming onsite. VR demonstration technique appeared in research projects like the iGibson home robot simulator, an initiative that I participated in at Stanford: Behind-the-scene video by Litian Liang

Jim Fan

124,783 Aufrufe • vor 2 Jahren