ๆญฃๅœจๅŠ ่ฝฝ่ง†้ข‘...

่ง†้ข‘ๅŠ ่ฝฝๅคฑ่ดฅ

Introduce Open-๐“๐ž๐ฅ๐ž๐•๐ข๐ฌ๐ข๐จ๐ง๐Ÿค–: โฃ We need an intuitive and remote teleoperation interface to collect more robot data. ๐“๐ž๐ฅ๐ž๐•๐ข๐ฌ๐ข๐จ๐ง lets you immersively operate a robot even if you are 3000 miles away, like in the movie ๐˜ˆ๐˜ท๐˜ข๐˜ต๐˜ข๐˜ณ. Open-sourced!

329,588 ๆฌก่ง‚็œ‹ โ€ข 2 ๅนดๅ‰ โ€ขvia X (Twitter)

10 ๆก่ฏ„่ฎบ

Xuxin Cheng ็š„ๅคดๅƒ
Xuxin Cheng2 ๅนดๅ‰

Real-time stereo video streaming provides spatial/depth understanding, so the operator is confident about the objectsโ€™ locations. This enables fine manipulation with challenging objects such as transparent boxes.

Xuxin Cheng ็š„ๅคดๅƒ
Xuxin Cheng2 ๅนดๅ‰

An active neck plus IK and retargeting enables intuitive perception and actuation for the operator. The operator just needs to look at the points of interest intuitively and the robot will follow the same head movements.

Xuxin Cheng ็š„ๅคดๅƒ
Xuxin Cheng2 ๅนดๅ‰

The system is easily accessible from any device with a web browser (Vision Pro, Quest, even mac, iPad, iPhoneโ€ฆ).

Xuxin Cheng ็š„ๅคดๅƒ
Xuxin Cheng2 ๅนดๅ‰

Not possible without the joint efforts of @Jialong_LI_UIM @AaronYANG2000 @EpisodeYang @xiaolonw Try now even if you donโ€™t have a VR device! More videos, code, hardware, and dataset at:

Huazhe Harry Xu ็š„ๅคดๅƒ
Huazhe Harry Xu2 ๅนดๅ‰

This is useful๏ผ

Xuxin Cheng ็š„ๅคดๅƒ
Xuxin Cheng2 ๅนดๅ‰

Thanks Huazhe!

Quanting Xie ็š„ๅคดๅƒ
Quanting Xie2 ๅนดๅ‰

This is awesome, congrats Xuxin!

Xuxin Cheng ็š„ๅคดๅƒ
Xuxin Cheng2 ๅนดๅ‰

Thanks Quanting!

BensenHsu ็š„ๅคดๅƒ
BensenHsu2 ๅนดๅ‰

The experiments show that the proposed system outperforms baseline models in terms of task success rates and completion times. The use of stereo video input is found to be crucial for the operator's spatial understanding and task performance. full paper:

Stefanos Charalambous ็š„ๅคดๅƒ
Stefanos Charalambous2 ๅนดๅ‰

@vateseif @arbwes

็›ธๅ…ณ่ง†้ข‘

We might be solving the wrong problem in robotics. Thatโ€™s what this makes clear. UMI โ†’ Universal Manipulation Interface A simple $400 gripper that lets you teach robots by demonstration. You hold it like a tool. Show the task. The robot learns. No teleoperation. No expensive hardware. No robot-specific data. Stanford open-sourced everything โ†’ hardware, code, datasets. What stands out to me is the bottleneck. Not algorithms. Data. Teleoperation โ†’ ~35 demos/hour UMI โ†’ ~111 demos/hour And the data transfers across robots โ†’ UR5, Franka, others. The design is surprisingly practical: โ†’ GoPro fisheye lens (155ยฐ FOV) + mirrors for depth โ†’ SLAM + IMU for precise 6DoF tracking โ†’ latency matching for dynamic tasks โ†’ diffusion policies for multimodal actions Then it scales. Cheng Chi takes this further with Sunday Robotics (with Tony Zhao). A $200 glove โ†’ deployed in 500+ homes โ†’ ~10 million real-world interactions. Not lab data. Real human behavior. Their robot learns dishes, laundry, espresso โ†’ with zero robot-specific data. This is where the shift becomes obvious. From training robots in controlled environments โ†’ to learning directly from humans at scale So hereโ€™s the real question: Will robotics be unlocked by better modelsโ€ฆ or by unlocking data? #ArtificialIntelligence #Robotics #AI #Innovation #FutureOfWork

Pascal Bornet

186,670 ๆฌก่ง‚็œ‹ โ€ข 4 ไธชๆœˆๅ‰