Loading video...

Video Failed to Load

Go Home

šŸ¤– What if a humanoid robot could make a hamburger from raw ingredients—all the way to your plate? šŸ”„ Excited to announce ViTacFormer: our new pipeline for next-level dexterous manipulation with active vision + high-resolution touch. šŸŽÆ For the first time ever, we demonstrate ~2.5 minutes of continuous, autonomous...

96,048 views • 1 year ago •via X (Twitter)

11 Comments

Haoran Geng's profile picture
Haoran Geng1 year ago

Our hardware setup features an active vision system, a bi-manual robot arm, and two high-DoF dexterous hands, SharpaWave, equipped with high-resolution tactile sensors. To enable rich data collection, we use a teleoperation system with a precision exoskeleton for fine-grained arm and finger control, plus a VR interface to control the camera and guide demonstrations.

Haoran Geng's profile picture
Haoran Geng1 year ago

ViTacFormer is a unified visuo-tactile framework for dexterous manipulation. At its core is a cross-modal representation that fuses visual and tactile inputs via cross-attention layers. Our key insight: predicting future tactile states is more informative than simply perceiving current ones. We validate this both conceptually and empirically. To achieve this, we introduce a tactile prediction head that encourages the shared latent space to capture touch dynamics. The model then auto-regressively leverages predicted tactile signals to generate precise, long-horizon actions.

Haoran Geng's profile picture
Haoran Geng1 year ago

We benchmarked our proposed pipeline against strong baselines across four challenging manipulation tasks. ViTacFormer significantly outperforms existing methods—achieving around 50% improvement in success rate.

Haoran Geng's profile picture
Haoran Geng1 year ago

Our ablation study highlights the contribution of each component, confirming their importance to the overall pipeline’s performance.

Haoran Geng's profile picture
Haoran Geng1 year ago

Here we also show the failure mode from baseline models, highlighting the importance of our tactile signal input, touch processing module, and cross-modal representation learning.

Haoran Geng's profile picture
Haoran Geng1 year ago

We then explored the full capabilities of our system—and found it can handle super long-horizon tasks end-to-end. šŸ” Sit back and enjoy the hamburger-making policy in action!

Haoran Geng's profile picture
Haoran Geng1 year ago

This work wouldn’t have been possible without the incredible support from great collaborators @cosm_para13983 and @KaifengZhang4, and my amazing advisors @JitendraMalikCV and @pabbeel. Thank you all! šŸ™ Code is fully released; check out our: Homepage: Paper link: Github:

Annie Lane's profile picture
Annie Lane1 year ago

AI isn’t one trade—it’s a layered ecosystem. Chips led, apps followed… what’s next?ā€ 🧠 We mapped the shift across AI subsectors

RobotSoul's profile picture
RobotSoul1 year ago

Awesome!

Jul3n3's profile picture
Jul3n31 year ago

It could significantly alter job roles and industry standards in food service.

Bianca Cortexfield's profile picture
Bianca Cortexfield1 year ago

Achieving full automation in food preparation presents complex challenges but could optimize consistency and hygiene.

Related Videos

X-Humanoid just officially dropped Embodied Tien Kung 3.0, A universal platform designed to be way more open and developer-friendly. šŸ¤– Built on their Wise Kaiwu AI platform, this next-gen humanoid is all about slashing development costs. It’s a fully interoperable ecosystem that supports everything from tactile interaction to high-dynamic motion control at a full humanoid scale. āž¤ Radical Openness: X-Humanoid is open-sourcing the full stack—robot body, motion control, VLM/VLA models, and the RoboMIND dataset. It fully supports ROS2, MQTT, and TCP/IP, so developers can customize use cases without re-engineering the basics. āž¤ High-Performance Hardware: With high-torque integrated joints, Tien Kung 3.0 can clear 1-meter (3.3ft) obstacles and handle dexterous moves like kneeling and bending. It hits millimeter-level precision, making it a solid fit for industrial-grade tasks. āž¤ True Autonomy: The bot runs a continuous perception-decision-execution loop. It uses world models to break down complex language commands and VLA models for real-time obstacle avoidance and navigation. āž¤ Scalable Collaboration: The platform moves beyond single-unit tasks to support multi-robot collaboration with autonomous scheduling. It’s built to move embodied AI from the lab straight into real-world commercial and industrial environments. Source: X-Humanoid #Humanoid #OpenSource #Robotics #EmbodiedAI #PhysicalAI #Automation #XHumanoid #TienKung #WiseKaiwu

RoboHubšŸ¤–

49,440 views • 7 months ago

We believe we’re the first robotics company to demonstrate a robot peeling an apple with dual dexterous human-like hands. This breakthrough closes a key gap in robotics, achieving bimanual, contact-rich manipulation and moving far beyond the limits of simple grippers. šŸ§µā†“ Today’s AI models (VLMs) are excellent at perception but struggle with action. Controlling high-degree-of-freedom hands for tasks like this is incredibly complex, and precise finger-level teleoperation is nearly impossible for humans. Our first step was a shared-autonomy system: rather than controlling every finger, the operator triggers pre-learned skills like a ā€œrotate apple or tennis ballā€ primitive via a keyboard press or pedal. This makes scalable data collection and RL training possible. How does the AI manage this? We created "MoDE-VLA" (Mixture of Dexterous Experts). It fuses vision, language, force, and touch data by using a team of specialist "experts," making control in high-dimensional spaces stable and effective. The combination of these two innovations allows for seamless, contact-rich manipulation. The human provides high-level guidance, and the robot executes the complex in-hand coordination required. This work paves the way for robots that can safely handle delicate tasks in human environments. Want the full technical details? šŸ“„ Read the full research paper: Visit us at NVIDIA GTC Booth #1838, Hall 3 to learn more! #Robotics #AI #DexterousManipulation #VLA #NVIDIAGTC Nancy VillicaƱa NVIDIA GTC

Sharpa

20,693 views • 7 months ago