Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

1/🧠Humans are the best robot data source — but video alone misses one thing: force. 2/🙁Tactile gloves capture force — but they're costly and block the real touch manipulation depends on. 3/💪Maybe the future of touch lives on your wrist: surface EMG reads the muscles that cause force —...

52,379 Aufrufe • vor 2 Monaten •via X (Twitter)

20 Kommentare

Profilbild von Zhi (Leo) Wang
Zhi (Leo) Wangvor 2 Monaten

「🔥INSIGHTS - 1/4 - sEMG Band Hardware Design」 🔥Muscle is the most upstream signal of human manipulation — muscles fire before a finger even moves, and they encode how hard each finger will press, even the ones a camera never sees. So we tap it at the source and built our own muscle-aware sEMG band: 1️⃣Electrodes over the exact forearm muscles that drive each finger — not evenly spaced 2️⃣sEMG + IMU in one wrist unit 3️⃣Custom fingertip force sensors — transparent gel + cables hidden palm-side, so we get ground-truth force without changing how the hand feels 4️⃣Fully DIY & extendable: 8 → 16 channels for denser signal Same electrode count, lower error — just by placing them on the right muscles. 🧵 2/n

Profilbild von Zhi (Leo) Wang
Zhi (Leo) Wangvor 2 Monaten

「🔥INSIGHTS - 2/4 - 10-Hours Dataset with EMG」 🔥Tactile data won't scale through gloves or lab rigs. It scales the way EMG does — you just wear a band and live your life. EMG is the most collectable touch signal we have: 1️⃣ Effortless & imperceptible — no gloves, no rigs, you barely feel it's on 2️⃣In-the-wild — collect anytime, anywhere, hands fully free 3️⃣Low-cost — vs >$10k lab force rigs 4️⃣Same form factor as egocentric video — one wearable, many modalities So we collected 10 hours of everyday forceful manipulation — synced sEMG + IMU + fingertip force + RGB, across many tasks and many grasp types. 🧵 3/n

Profilbild von Zhi (Leo) Wang
Zhi (Leo) Wangvor 2 Monaten

「🔥INSIGHTS - 3/4 - EMG2Force」 🔥Vision can barely tell you whether a finger is in contact — let alone how hard. But muscle is where force is born: it fires before motion, before contact is even visible. From sEMG + IMU, ForceBand predicts continuous, per-finger force for all 5 fingertips — in real Newtons: 1️⃣A transformer fuses time + frequency-domain (spectrogram) signals — trained on the 10h dataset above 2️⃣Per-finger & continuous — not a binary grasp flag 3️⃣Halves the error of the best vision force estimator (MAE 0.80N vs 1.61N) 4️⃣Wins biggest on fingers vision can't see — pinky 0.26N vs 0.73N 5️⃣Tracks force as it changes — object-specific & time-varying EMG sees the forces that vision simply cannot. 🧵 4/n

Profilbild von Zhi (Leo) Wang
Zhi (Leo) Wangvor 2 Monaten

「🔥INSIGHTS - 4/4 - Fully Policy: Vision+Force」 🔥EMG isn't a replacement for vision — it's the missing half. ForceBand = emg2force + HumanEgo ( our framework for learning robot policies from human video alone (robot-data-free). 1️⃣ HumanEgo turns raw egocentric video into an embodiment-agnostic observation: hand + object tracking, arm inpainting, a virtual gripper rendered in. 2️⃣ emg2force adds the missing modality — predicted fingertip force, time-aligned to every frame. 3️⃣ A flow-matching policy reads force on the observation and predicts it on the action: a = [pos, rot, grasp, force]. 4️⃣ At deploy, a PD controller closes the loop on the gripper's force sensor. Vision-only HumanEgo → force-aware ForceBand. Same recipe, now with force. Now it can do the forceful tasks vision-only policies (like HumanEgo) simply can't: ✦ Squeeze toothpaste, mustard, ketchup — just hard enough ✦ Place soft bread & eggs without crushing them ✦ Modulate grip per object, in real time Force-aware: 8–10/10. Vision-only / binary gripper: ~0. Same human video — just +force. Vision tells the robot where. Force tells it how hard. You need both! 🧵 5/n

Profilbild von Zhi (Leo) Wang
Zhi (Leo) Wangvor 2 Monaten

「🤖 ForceBand → Robot, in 3 Steps」 1️⃣ Calibrate — 15 min of random play with the band + fingertip force sensors. Learns your EMG → force map. 2️⃣ Collect — sensors off. Just the band on a bare hand, recording everyday demos in the wild — force predicted live by emg2force. 3️⃣ Deploy — train a force-aware policy, run it zero-shot on the robot. It predicts object-specific force and modulates grip in real time. No teleop. No robot data. The force sensors come off after 15 min — the rest is just a band, human video, and force. 🧵 6/n

Profilbild von Zhi (Leo) Wang
Zhi (Leo) Wangvor 2 Monaten

「🌍 What can a force-aware policy actually do?」 In the real world, the robot applies the right force for each object: 1️⃣Squeeze condiments — mustard 22.1N, BBQ 17–18N, facewash 15.3N, toothpaste ~10N — dispense without bursting the bottle 2️⃣Pick & place delicate items — chips at 1.7N, soft bread & eggs, no crushing 3️⃣Twist & lift bottle caps One policy, forces from ~2N to 22N, set per object in real time. A binary / continuous-gripper policy can only do one thing — clamp. 🧵 7/n

Profilbild von Zhi (Leo) Wang
Zhi (Leo) Wangvor 2 Monaten

「🛡️ Does it hold up out of distribution?」 Yes — because force comes from your muscle, not from pixels. 1️⃣New objects & sizes — trained on regular objects, generalizes to unseen, smaller & larger ones 2️⃣Visual shift — occlusion, clutter, lighting, novel appearance don't break the force signal (vision-based force estimators do) 3️⃣New force profiles — handles magnitudes & dynamics it wasn't trained on Vision-based force degrades exactly when the scene gets hard. Muscle-based force doesn't. 🧵 8/n

Profilbild von Zhi (Leo) Wang
Zhi (Leo) Wangvor 2 Monaten

「🔥 EMG might not stop at force!」 The same wrist signal could decode 3D force, full hand motion — even under complete visual occlusion. That's exactly where human video breaks, and exactly what robot policies are missing. We believe cheap, comfortable EMG is the next platform for capturing human interaction: 👓 Egocentric glasses → what the hand sees ⌚️ Wristband → what the hand feels Low-cost, easy to wear, easy to use, scalable. We hope ForceBand sparks more thinking about EMG — and the EMG + Ego form factor — as a new way to capture human data. Vision gave robots the world's appearance — EMG can give them its forces! 🧵 9/n

Profilbild von Zhi (Leo) Wang
Zhi (Leo) Wangvor 2 Monaten

「🙏 Thank you!!!」 ForceBand is a joint effort across @amazon FAR Team, University of Maryland @UofMaryland , and Johns Hopkins University @JohnsHopkins Huge thanks to our incredible team: @BotaoUMD (Lead), @TX_Leo_Wang , Linna Kuang, Ishaan Ghosh Deepest thanks to the advisors: @JitendraMalikCV , @CFermuller , Tingfan Wu, @maojiayuan , @ruoshi_liu , @HaozhiQ , @YAloimonos 🌐 Website: 📄 Paper: 💻 Code: 🎥 Video: 🧵10/n

Profilbild von Alfred Cueva
Alfred Cuevavor 2 Monaten

Force aware policies is what get us closer to real deployment, nice work!

Profilbild von Zhi (Leo) Wang
Zhi (Leo) Wangvor 2 Monaten

Thank you, Alfred!

Profilbild von chuong nguyen
chuong nguyenvor 2 Monaten

Very nice idea

Profilbild von Angkul
Angkulvor 2 Monaten

This is huge

Profilbild von Michael Vasilkovsky
Michael Vasilkovskyvor 2 Monaten

Love this direction! Having inferred force feedback as part of egocentric pre-training would be a massive unlock

Profilbild von Nikhil Nakhate
Nikhil Nakhatevor 2 Monaten

This is so cool, it sounds unreal!

Profilbild von Jakie PLA
Jakie PLAvor 2 Monaten

Per-finger force from sEMG. 50% lower error than vision baselines. THIS is what robot manipulation learning has needed for YEARS. I need to try this.

Profilbild von csgm
csgmvor 2 Monaten

Incredible work! This is much better than teleop! I love the idea! Can't wait to see an actual large scale dataset! This + good retargeting + sharpa/wuji-like hand is going to be the future i'm sure!

Profilbild von Pushkar
Pushkarvor 2 Monaten

Great reading. What i'm trying to understand The calibration step still relying on fingertip force sensors. Do you think this is ultimately a data scaling problem (enough diverse sEMG + force pairs), or do you expect some form of calibration to always be necessary across users?

Profilbild von SEAR
SEARvor 2 Monaten

Wrist sensors are cool

Profilbild von Rachel T
Rachel Tvor 2 Monaten

Very promising. I think force is the missing variable in much of imitation learning, and a low-cost open-source way to capture it could materially improve robotic manipulation.

Ähnliche Videos

I was really impressed by the UMI gripper (Cheng Chi et al.), but a key limitation is that **force-related data wasn’t captured**: humans feel haptic feedback through the mechanical springs, but the robot couldn’t leverage that info, limiting the data’s value for fine-grained manipulation tasks. Led by my amazing students Yolanda Zhu and Binghao Huang, we designed a **portable visuo-tactile gripper** by integrating our dense, flexible tactile arrays with the UMI gripper to enable large-scale in-the-wild data collection. 🔗 We demonstrate **cross-modal representation learning** and **downstream policy learning** on tasks requiring in-hand state estimation (e.g., test tube reorientation) and fine-grained force sensing (e.g., pipette fluid transfer). Key takeaways: - Our flexible tactile arrays store the rich haptic information humans perceive as dense tactile signals. - Portability and robustness are key for in-the-wild data collection; our portable gripper is compact, lightweight, and durable. - Touch provides precise, robust measurements of in-hand object pose, invariant to lighting and viewpoint. - Cross-modal pretraining on large-scale in-the-wild data significantly improves policy robustness and sample efficiency (as shown many times before — and verified again here!). Also check out our previous investigations of dense, flexible tactile grids for understanding human-robot-environment interactions: - Dense tactile glove (Nature ’19): - 3D-ViTac (CoRL ’24):

Yunzhu Li

13,188 Aufrufe • vor 1 Jahr

Force-sensing fingers! 🧤 Stanford researchers just released UMI-FT, a handheld data collection platform that puts compact six-axis force/torque sensors on each finger, enabling finger-level wrench measurements alongside RGB, depth, and pose data. Many manipulation tasks require careful force modulation: too little force and the task fails, too much and you cause damage. But commercial force/torque sensors are expensive, bulky, and fragile, which has limited large-scale force-aware policy learning. UMI-FT changes the economics. The platform uses an iPhone for RGB vision, ultrawide RGB, depth, and pose via ARKit, with each finger sensorized using a CoinFT sensor to capture per-finger wrench information during manipulation. This multimodal data trains an adaptive compliance policy that predicts position targets, grasp force, and stiffness for execution on standard compliance controllers. The learned policy runs slowest and generates reference targets, while model-based compliance and force controllers provide delicate 6D compliance control and real-time force modulation. They tested on three contact-rich, force-sensitive tasks: whiteboard wiping (locate eraser, grasp, wipe until clean), skewering zucchini (grasp slice firmly, push onto stick until punctured), and lightbulb insertion (grasp bulb, align bayonet pin with socket slit, insert while overcoming spring force, rotate to light up). The results are clear. Policies without compliance struggle to modulate contact force and trigger safety faults from excessive force. Policies without force sensing fail to grasp unseen objects or resist reaction forces, causing slippage. Here's the project page: ~~ ♻️ Join the weekly robotics newsletter, and never miss any news →

Lukas Ziegler

12,868 Aufrufe • vor 7 Monaten

The sense of touch is the most criminally under-explored modality in robotics. Imagine doing sleight of hand wearing thick oven mitts. That's exactly how a robot feels today if it were alive. A magnetic piece snapping into place, a paper cup peeling out of a stack, a USB negotiating its way into the port - all invisible to the camera. Learning how to feel must be a full-stack co-designed effort. We are open-sourcing a principled methodology called "T-Rex": 1. Tactile as first-class citizen of the model. Our mixture-of-transformer runs two clocks asynchronously: a slow visuomotor expert plans the motion, and a fast tactile expert refines it in real time with high-frequency corrections at 4 "touch ticks" per vision tick. Forces change faster than frames arrive, so the architecture had to as well. 2. Open data. The largest tactile dataset ever released to our knowledge: a 50-hour (~5,500 episodes) high-quality, carefully synchronized robot play corpus, collected on SOTA tactile hand hardware with 22 degrees of freedom. Available today on HuggingFace! 3. Training recipe: T-Rex extends our prior work, EgoScale. Human egocentric videos for pretraining, a diverse dose of tactile robot play for mid-training. Our experiments show this bridges contact-free pretraining to contact-rich manipulation remarkably well. Pixels are cheap and everywhere, but they run out of steam at the moment of contact. Tactile will carry the last mile. The next scaling curve will be measured in hours of touch. T-Rex is a great collaboration between NVIDIA and Berkeley: 🧵

Jim Fan

173,941 Aufrufe • vor 1 Monat

🔥 JUST IN: Open-source robotics dataset from 100% real-world scenarios! 🤯 Chinese robotics company AGIBOT just released AGIBOT WORLD 2026, an open-source dataset systematically covering key embodied AI research directions. Built entirely from real-world environments: commercial spaces, and homes. Collected using AGIBOT G2 robots in free-form collection mode, providing structured, accurately annotated, high-quality data. Digital twin technology creates 1:1 scale replicas in simulation matching the real environments. Both real-world and simulation data are open-sourced. The AGIBOT G2 platform collects multiple data types simultaneously: RGB(D) cameras, tactile sensors, force sensors, LiDAR, IMU, and full-body joint states. Whole-body control coordinates arms, waist, and hands for complex tasks. First-person teleoperation lets operators control the robot from its perspective. The tasks covered are fine-grained manipulation, ultra-long-horizon tasks, spatial navigation, dual-arm coordination, and multi-agent/human-robot collaboration. The dataset includes error-recovery trajectories with annotations. Most datasets only show successful demonstrations. AGIBOT includes failures and how the robot recovers, teaching models how to handle mistakes. After collection, data is tested through policy training and real-robot deployment to ensure quality. Then processed through industrial quality control with multiple screening and cleaning rounds. Making it open-source accelerates embodied AI research by giving researchers access to high-quality real-world robot data at scale. 🇨🇳 Learn more here: ~~ ♻️ Join the weekly robotics newsletter, and never miss any news →

Lukas Ziegler

40,583 Aufrufe • vor 5 Monaten