Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Presenting research from Berkeley AI Research Humanoid Intelligence Center. DexTacWAM: A Visuo-Tactile World-Action Model for Dexterous Manipulation.🖐️🤖 Dexterous manipulation requires touch, yet multi-finger tactile data remain scarce and expensive to collect at scale. Through continual vision-to-touch learning, DexTacWAM adapts a pretrained video world model into a visuo-tactile world model...

109,600 Aufrufe • vor 6 Tagen •via X (Twitter)

11 Kommentare

Profilbild von Berkeley AI Research
Berkeley AI Researchvor 6 Tagen

Full results across six contact-rich tasks, with 20 real-robot trials per method per task and identical camera access for all methods. DexTacWAM achieves the highest score on every task, averaging 70.6 vs. 38.0 for RDP, the strongest baseline. The gap is largest where vision is least informative: on the tongs task, every baseline scores ≤10, while DexTacWAM reaches 60. Importantly, RDP and ViTacFormer already receive fingertip tactile input. The gain is not from touch alone, but from modeling how contact evolves.

Profilbild von Berkeley AI Research
Berkeley AI Researchvor 6 Tagen

Why vision-to-touch transfer? Large-scale human video is abundant and inexpensive to collect, while tactile data still depend heavily on physical interaction and robot teleoperation. DexTacWAM adapts the tactile encoder using only 4 hours of tactile interaction data, then directly extends a pretrained video world model to tactile prediction using approximately 100 demonstrations per task, without tactile midtraining of the video backbone.

Profilbild von Berkeley AI Research
Berkeley AI Researchvor 6 Tagen

Why predict touch rather than simply condition on it? Future tactile prediction provides a self-supervised objective for learning contact dynamics. At inference, a single world-model forward pass provides predictive visuo-tactile features directly to the action expert, without requiring fully denoised future visual or tactile observations.

Profilbild von Berkeley AI Research
Berkeley AI Researchvor 6 Tagen

How do we scale tactile world modeling across ten fingertips? A finger- and pose-aware tactile compressor maps: 10 fingertip streams → 2 hand-level latents while retaining 89.4% of pre-fusion contact recall. This enables 2.26× faster training and 1.29× faster inference.

Profilbild von Berkeley AI Research
Berkeley AI Researchvor 6 Tagen

The project website includes additional real-robot demonstrations, world-model predictions, tactile visualizations, ablations, and detailed explanations. 🌐

Profilbild von Patch
Patchvor 6 Tagen

The direct-touch comparison is the part that grabbed me. With the same encoder and policy, 74.7 vs. 26.6 is a striking gap between predicting contact and simply feeding touch to the policy.

Profilbild von The SI Therapist
The SI Therapistvor 6 Tagen

DexTacWAM uses touch to help robots handle objects well. This mix of sight and feel makes dexterous work smoother for AI hands.

Profilbild von Howard Alvaro
Howard Alvarovor 5 Tagen

This research addresses a critical gap in tactile data collection for manipulation. Can lead to significant advancements in robotics.

Profilbild von NAMAN RAJ
NAMAN RAJvor 6 Tagen

Sounds like we're about to give robots a sense of touch that's more than just a feel-good gesture 🤖💨

Profilbild von Julie Maginnis
Julie Maginnisvor 5 Tagen

Hopeful for beneficial applications of this technology to assist those in need of assistance following injury from paralysis, trauma & neuropathy. Excellent work!

Profilbild von Hershal Rao
Hershal Raovor 6 Tagen

finally, robots can learn the art of touching grass

Ähnliche Videos

The sense of touch is the most criminally under-explored modality in robotics. Imagine doing sleight of hand wearing thick oven mitts. That's exactly how a robot feels today if it were alive. A magnetic piece snapping into place, a paper cup peeling out of a stack, a USB negotiating its way into the port - all invisible to the camera. Learning how to feel must be a full-stack co-designed effort. We are open-sourcing a principled methodology called "T-Rex": 1. Tactile as first-class citizen of the model. Our mixture-of-transformer runs two clocks asynchronously: a slow visuomotor expert plans the motion, and a fast tactile expert refines it in real time with high-frequency corrections at 4 "touch ticks" per vision tick. Forces change faster than frames arrive, so the architecture had to as well. 2. Open data. The largest tactile dataset ever released to our knowledge: a 50-hour (~5,500 episodes) high-quality, carefully synchronized robot play corpus, collected on SOTA tactile hand hardware with 22 degrees of freedom. Available today on HuggingFace! 3. Training recipe: T-Rex extends our prior work, EgoScale. Human egocentric videos for pretraining, a diverse dose of tactile robot play for mid-training. Our experiments show this bridges contact-free pretraining to contact-rich manipulation remarkably well. Pixels are cheap and everywhere, but they run out of steam at the moment of contact. Tactile will carry the last mile. The next scaling curve will be measured in hours of touch. T-Rex is a great collaboration between NVIDIA and Berkeley: 🧵

Jim Fan

178,632 Aufrufe • vor 1 Monat

Sharpa Robotics just dropped a new hand video and the level keeps going up. This is the Sharpa Wave running WM Craftnet on a human scale fivefinger hand with 22 active DoF. The policy combines wrist depth, tactile sensing, proprioception and previous actions. The hand can rotate different objects in-hand, recover after external pushes and continue manipulating objects it was never trained on. The numbers are strong. 175/200 successful real-world rotation trials across 20 objects. A world-model prior trained on 9 objects was transferred to 49 new objects. Fall rate went from 6% to 0.3%. The Wave hardware itself has 22 actuators, up to 20 N fingertip force, 240×240 tactile sensing at up to 180 fps and 0.02 N pressure sensitivity. What caught my attention is the recovery behavior. The fingers keep changing contact points after the object slips or gets pushed instead of replaying the same finger motion. That is the kind of dexterity I want to see more of in robotic hands. According to Sharpa’s current specifications: • DTA tactile sensors on the fingers with a resolution of up to 240 × 240 • Pressure detection • Slip detection • Force change detection • Contact point localization • 6-axis force and torque measurement: Fx, Fy, Fz, Mx, My, Mz • Tactile sensing at up to 180 fps • 20 ms reported latency • Force detection range from 0 to 30 N • Maximum sensor load of 50 N • Sharpa also describes a miniature camera integrated into each fingertip for visuo-tactile sensing.

Techniahqrobot | humanoid robots

13,500 Aufrufe • vor 16 Tagen

Meta just open-sourced its dexterity stack! 🪬 Most physics simulators were built for things that move through space. Walking robots, drones, cars. Contact is the part they approximate worst, and obviously dexterous manipulation is nothing but contact. Project SuperDex, from Meta Reality Labs Research, is built the other way around, a contact-first physics engine with the whole platform stacked on top of it. The cool part is that it's on GitHub. The engine runs one solver across rigid bodies, soft bodies, rods and tendons, shells and cloth, in the same model. → Non-convex collision with accurate contact force distributions, so a multi-finger grasp gets simulated rather than approximated → Tactile sensors and soft contact as first-class primitives → Numerical stability without the tight time-step limits explicit solvers force on you → Constraint-aware inverse kinematics running on the same optimization core as the forward dynamics Then the data layer. Put on a Quest 3, teleoperate the simulated hand with haptic feedback, and generate demonstration datasets without touching real hardware. They show a shape-sorting policy trained entirely in simulation and deployed zero-shot on a real robotic hand. Robot hands are getting good. Data for them isn't that fast. Meta is betting the cheapest way to collect contact-rich demonstrations is a headset people already own, pointed at a simulator instead of a game. 🔗 Here's the project page: ~~ ♻️ Join the weekly robotics newsletter, and never miss any news →

Lukas Ziegler

60,709 Aufrufe • vor 1 Monat

A policy that teaches robot hands to touch things the way humans do... not just grab and move, but feel and adjust in real time. Robot manipulation research often stops at picking up objects and placing them. CGP goes further: it handles tasks like opening jars, flipping objects in-hand, wiping dishes, and grasping fragile eggs, the kind of dexterous, contact-rich skills that require constant micro-adjustments based on what the fingers are actually feeling. The robot doesn't just see what it's doing; it predicts what contact should feel like at each step, then checks whether reality matches the prediction. If a finger is slipping, the policy knows before the object drops. Works on real robot hands (both 4-finger and 5-finger designs) with tactile sensors embedded in the fingertips Robust to visual distractions! The robot keeps flipping a box correctly even when the camera view is disrupted, because it's grounding decisions in touch, not just vision. Baseline policies without contact grounding fail in predictable ways: slipping mid-task, incomplete motions, loss of grasp, CGP avoids these This is a meaningful step toward robots that can handle the physical world with the kind of reliable, adaptive grip that humans take for granted. Relevant for manufacturing, logistics, assistive robotics, and anywhere fragile or irregular objects need to be handled carefully. Published at RSS 2026, developed with Meta Reality Labs Research. Thanks for sharing, Zhengtong Xu / Zhengtong Xu ——- Weekly robotics and AI insights. Subscribe free:

Ilir Aliu

12,860 Aufrufe • vor 4 Monaten

Chinese robotics company Astribot released their latest World-Action Model (WAM), Lumo-2. Technical breakdown: - based on a frozen 🥶 Qwen-3.5 4B VLM - trained in 3 progressive stages: 1. Action is aligned with latent world dynamics (an abstract representation of action). Real-world actions are anchored to physical constraints, while the latent space is guided to focus on motion-relevant changes. This bidirectional relationship makes the model physically grounded -> critical for a world model. 2. Action is aligned with vision and language. Reusing the vision backbone and action encoder from the frozen VLM, the authors add a custom vocabulary (for new actions), a semantic module, an action decoder, and an action projector. This aligns the (new) action representations with the (existing) vision-language semantic space. Most importantly: it builds a direct mapping from natural-language instructions to motor execution. 3. End-to-end training on language, video, and robot data. Only the new modules (everything outside the frozen backbone) are trained end-to-end across temporal reasoning, physical understanding, long-horizon, and dexterous manipulation. At the end of the day, Lumo-2 is not the best on benchmarks, but that's not the point. What's genuinely new: - a way to combine latent world modeling and action generation through progressive alignment - a physically-grounded latent dynamics space - it lifts performance on unseen objects using un-annotated human egocentric video + Vision Pro captures, no special transfer algorithm needed Why it matters: - the whole model is thin trainable adapters (semantic module, action decoder/projector) on a frozen 4B backbone (cheap) - that scale is suited for real-time embedded inference (~2.71× decode speedup, no accuracy loss) - its real moat is long-horizon execution, where the added temporal memory pays off far more than on any other task As a result, this robot can now make your latte (5x sped up video):

Léo

32,513 Aufrufe • vor 2 Monaten