正在加载视频...

视频加载失败

🤖 What if a humanoid robot could make a hamburger from raw ingredients—all the way to your plate? 🔥 Excited to announce ViTacFormer: our new pipeline for next-level dexterous manipulation with active vision + high-resolution touch. 🎯 For the first time ever, we demonstrate ~2.5 minutes of continuous, autonomous...

95,703 次观看 • 1 年前 •via X (Twitter)

11 条评论

Haoran Geng 的头像
Haoran Geng1 年前

Our hardware setup features an active vision system, a bi-manual robot arm, and two high-DoF dexterous hands, SharpaWave, equipped with high-resolution tactile sensors. To enable rich data collection, we use a teleoperation system with a precision exoskeleton for fine-grained arm and finger control, plus a VR interface to control the camera and guide demonstrations.

Haoran Geng 的头像
Haoran Geng1 年前

ViTacFormer is a unified visuo-tactile framework for dexterous manipulation. At its core is a cross-modal representation that fuses visual and tactile inputs via cross-attention layers. Our key insight: predicting future tactile states is more informative than simply perceiving current ones. We validate this both conceptually and empirically. To achieve this, we introduce a tactile prediction head that encourages the shared latent space to capture touch dynamics. The model then auto-regressively leverages predicted tactile signals to generate precise, long-horizon actions.

Haoran Geng 的头像
Haoran Geng1 年前

We benchmarked our proposed pipeline against strong baselines across four challenging manipulation tasks. ViTacFormer significantly outperforms existing methods—achieving around 50% improvement in success rate.

Haoran Geng 的头像
Haoran Geng1 年前

Our ablation study highlights the contribution of each component, confirming their importance to the overall pipeline’s performance.

Haoran Geng 的头像
Haoran Geng1 年前

Here we also show the failure mode from baseline models, highlighting the importance of our tactile signal input, touch processing module, and cross-modal representation learning.

Haoran Geng 的头像
Haoran Geng1 年前

We then explored the full capabilities of our system—and found it can handle super long-horizon tasks end-to-end. 🍔 Sit back and enjoy the hamburger-making policy in action!

Haoran Geng 的头像
Haoran Geng1 年前

This work wouldn’t have been possible without the incredible support from great collaborators @cosm_para13983 and @KaifengZhang4, and my amazing advisors @JitendraMalikCV and @pabbeel. Thank you all! 🙏 Code is fully released; check out our: Homepage: Paper link: Github:

Annie Lane 的头像
Annie Lane1 年前

AI isn’t one trade—it’s a layered ecosystem. Chips led, apps followed… what’s next?” 🧠 We mapped the shift across AI subsectors

RobotSoul 的头像
RobotSoul1 年前

Awesome!

Jul3n3 的头像
Jul3n31 年前

It could significantly alter job roles and industry standards in food service.

Bianca Cortexfield 的头像
Bianca Cortexfield1 年前

Achieving full automation in food preparation presents complex challenges but could optimize consistency and hygiene.

相关视频

X-Humanoid just officially dropped Embodied Tien Kung 3.0, A universal platform designed to be way more open and developer-friendly. 🤖 Built on their Wise Kaiwu AI platform, this next-gen humanoid is all about slashing development costs. It’s a fully interoperable ecosystem that supports everything from tactile interaction to high-dynamic motion control at a full humanoid scale. ➤ Radical Openness: X-Humanoid is open-sourcing the full stack—robot body, motion control, VLM/VLA models, and the RoboMIND dataset. It fully supports ROS2, MQTT, and TCP/IP, so developers can customize use cases without re-engineering the basics. ➤ High-Performance Hardware: With high-torque integrated joints, Tien Kung 3.0 can clear 1-meter (3.3ft) obstacles and handle dexterous moves like kneeling and bending. It hits millimeter-level precision, making it a solid fit for industrial-grade tasks. ➤ True Autonomy: The bot runs a continuous perception-decision-execution loop. It uses world models to break down complex language commands and VLA models for real-time obstacle avoidance and navigation. ➤ Scalable Collaboration: The platform moves beyond single-unit tasks to support multi-robot collaboration with autonomous scheduling. It’s built to move embodied AI from the lab straight into real-world commercial and industrial environments. Source: X-Humanoid #Humanoid #OpenSource #Robotics #EmbodiedAI #PhysicalAI #Automation #XHumanoid #TienKung #WiseKaiwu

RoboHub🤖

49,440 次观看 • 6 个月前

We believe we’re the first robotics company to demonstrate a robot peeling an apple with dual dexterous human-like hands. This breakthrough closes a key gap in robotics, achieving bimanual, contact-rich manipulation and moving far beyond the limits of simple grippers. 🧵↓ Today’s AI models (VLMs) are excellent at perception but struggle with action. Controlling high-degree-of-freedom hands for tasks like this is incredibly complex, and precise finger-level teleoperation is nearly impossible for humans. Our first step was a shared-autonomy system: rather than controlling every finger, the operator triggers pre-learned skills like a “rotate apple or tennis ball” primitive via a keyboard press or pedal. This makes scalable data collection and RL training possible. How does the AI manage this? We created "MoDE-VLA" (Mixture of Dexterous Experts). It fuses vision, language, force, and touch data by using a team of specialist "experts," making control in high-dimensional spaces stable and effective. The combination of these two innovations allows for seamless, contact-rich manipulation. The human provides high-level guidance, and the robot executes the complex in-hand coordination required. This work paves the way for robots that can safely handle delicate tasks in human environments. Want the full technical details? 📄 Read the full research paper: Visit us at NVIDIA GTC Booth #1838, Hall 3 to learn more! #Robotics #AI #DexterousManipulation #VLA #NVIDIAGTC Nancy Villicaña NVIDIA GTC

Sharpa

20,429 次观看 • 5 个月前