正在加载视频...
视频加载失败
Presenting research from Berkeley AI Research Humanoid Intelligence Center. DexTacWAM: A Visuo-Tactile World-Action Model for Dexterous Manipulation.🖐️🤖 Dexterous manipulation requires touch, yet multi-finger tactile data remain scarce and expensive to collect at scale. Through continual vision-to-touch learning, DexTacWAM adapts a pretrained video world model into a visuo-tactile world model... show more
11 条评论

Full results across six contact-rich tasks, with 20 real-robot trials per method per task and identical camera access for all methods. DexTacWAM achieves the highest score on every task, averaging 70.6 vs. 38.0 for RDP, the strongest baseline. The gap is largest where vision is least informative: on the tongs task, every baseline scores ≤10, while DexTacWAM reaches 60. Importantly, RDP and ViTacFormer already receive fingertip tactile input. The gain is not from touch alone, but from modeling how contact evolves.

Why vision-to-touch transfer? Large-scale human video is abundant and inexpensive to collect, while tactile data still depend heavily on physical interaction and robot teleoperation. DexTacWAM adapts the tactile encoder using only 4 hours of tactile interaction data, then directly extends a pretrained video world model to tactile prediction using approximately 100 demonstrations per task, without tactile midtraining of the video backbone.

Why predict touch rather than simply condition on it? Future tactile prediction provides a self-supervised objective for learning contact dynamics. At inference, a single world-model forward pass provides predictive visuo-tactile features directly to the action expert, without requiring fully denoised future visual or tactile observations.

How do we scale tactile world modeling across ten fingertips? A finger- and pose-aware tactile compressor maps: 10 fingertip streams → 2 hand-level latents while retaining 89.4% of pre-fusion contact recall. This enables 2.26× faster training and 1.29× faster inference.

The project website includes additional real-robot demonstrations, world-model predictions, tactile visualizations, ablations, and detailed explanations. 🌐

The direct-touch comparison is the part that grabbed me. With the same encoder and policy, 74.7 vs. 26.6 is a striking gap between predicting contact and simply feeding touch to the policy.

DexTacWAM uses touch to help robots handle objects well. This mix of sight and feel makes dexterous work smoother for AI hands.

This research addresses a critical gap in tactile data collection for manipulation. Can lead to significant advancements in robotics.

Sounds like we're about to give robots a sense of touch that's more than just a feel-good gesture 🤖💨

Hopeful for beneficial applications of this technology to assist those in need of assistance following injury from paralysis, trauma & neuropathy. Excellent work!

finally, robots can learn the art of touching grass

