Video wird geladen...
Video konnte nicht geladen werden
1/🧠Humans are the best robot data source — but video alone misses one thing: force. 2/🙁Tactile gloves capture force — but they're costly and block the real touch manipulation depends on. 3/💪Maybe the future of touch lives on your wrist: surface EMG reads the muscles that cause force —... show more
52,379 Aufrufe • vor 2 Monaten •via X (Twitter)
20 Kommentare

「🔥INSIGHTS - 1/4 - sEMG Band Hardware Design」 🔥Muscle is the most upstream signal of human manipulation — muscles fire before a finger even moves, and they encode how hard each finger will press, even the ones a camera never sees. So we tap it at the source and built our own muscle-aware sEMG band: 1️⃣Electrodes over the exact forearm muscles that drive each finger — not evenly spaced 2️⃣sEMG + IMU in one wrist unit 3️⃣Custom fingertip force sensors — transparent gel + cables hidden palm-side, so we get ground-truth force without changing how the hand feels 4️⃣Fully DIY & extendable: 8 → 16 channels for denser signal Same electrode count, lower error — just by placing them on the right muscles. 🧵 2/n

「🔥INSIGHTS - 2/4 - 10-Hours Dataset with EMG」 🔥Tactile data won't scale through gloves or lab rigs. It scales the way EMG does — you just wear a band and live your life. EMG is the most collectable touch signal we have: 1️⃣ Effortless & imperceptible — no gloves, no rigs, you barely feel it's on 2️⃣In-the-wild — collect anytime, anywhere, hands fully free 3️⃣Low-cost — vs >$10k lab force rigs 4️⃣Same form factor as egocentric video — one wearable, many modalities So we collected 10 hours of everyday forceful manipulation — synced sEMG + IMU + fingertip force + RGB, across many tasks and many grasp types. 🧵 3/n

「🔥INSIGHTS - 3/4 - EMG2Force」 🔥Vision can barely tell you whether a finger is in contact — let alone how hard. But muscle is where force is born: it fires before motion, before contact is even visible. From sEMG + IMU, ForceBand predicts continuous, per-finger force for all 5 fingertips — in real Newtons: 1️⃣A transformer fuses time + frequency-domain (spectrogram) signals — trained on the 10h dataset above 2️⃣Per-finger & continuous — not a binary grasp flag 3️⃣Halves the error of the best vision force estimator (MAE 0.80N vs 1.61N) 4️⃣Wins biggest on fingers vision can't see — pinky 0.26N vs 0.73N 5️⃣Tracks force as it changes — object-specific & time-varying EMG sees the forces that vision simply cannot. 🧵 4/n

「🔥INSIGHTS - 4/4 - Fully Policy: Vision+Force」 🔥EMG isn't a replacement for vision — it's the missing half. ForceBand = emg2force + HumanEgo ( our framework for learning robot policies from human video alone (robot-data-free). 1️⃣ HumanEgo turns raw egocentric video into an embodiment-agnostic observation: hand + object tracking, arm inpainting, a virtual gripper rendered in. 2️⃣ emg2force adds the missing modality — predicted fingertip force, time-aligned to every frame. 3️⃣ A flow-matching policy reads force on the observation and predicts it on the action: a = [pos, rot, grasp, force]. 4️⃣ At deploy, a PD controller closes the loop on the gripper's force sensor. Vision-only HumanEgo → force-aware ForceBand. Same recipe, now with force. Now it can do the forceful tasks vision-only policies (like HumanEgo) simply can't: ✦ Squeeze toothpaste, mustard, ketchup — just hard enough ✦ Place soft bread & eggs without crushing them ✦ Modulate grip per object, in real time Force-aware: 8–10/10. Vision-only / binary gripper: ~0. Same human video — just +force. Vision tells the robot where. Force tells it how hard. You need both! 🧵 5/n

「🤖 ForceBand → Robot, in 3 Steps」 1️⃣ Calibrate — 15 min of random play with the band + fingertip force sensors. Learns your EMG → force map. 2️⃣ Collect — sensors off. Just the band on a bare hand, recording everyday demos in the wild — force predicted live by emg2force. 3️⃣ Deploy — train a force-aware policy, run it zero-shot on the robot. It predicts object-specific force and modulates grip in real time. No teleop. No robot data. The force sensors come off after 15 min — the rest is just a band, human video, and force. 🧵 6/n

「🌍 What can a force-aware policy actually do?」 In the real world, the robot applies the right force for each object: 1️⃣Squeeze condiments — mustard 22.1N, BBQ 17–18N, facewash 15.3N, toothpaste ~10N — dispense without bursting the bottle 2️⃣Pick & place delicate items — chips at 1.7N, soft bread & eggs, no crushing 3️⃣Twist & lift bottle caps One policy, forces from ~2N to 22N, set per object in real time. A binary / continuous-gripper policy can only do one thing — clamp. 🧵 7/n

「🛡️ Does it hold up out of distribution?」 Yes — because force comes from your muscle, not from pixels. 1️⃣New objects & sizes — trained on regular objects, generalizes to unseen, smaller & larger ones 2️⃣Visual shift — occlusion, clutter, lighting, novel appearance don't break the force signal (vision-based force estimators do) 3️⃣New force profiles — handles magnitudes & dynamics it wasn't trained on Vision-based force degrades exactly when the scene gets hard. Muscle-based force doesn't. 🧵 8/n

「🔥 EMG might not stop at force!」 The same wrist signal could decode 3D force, full hand motion — even under complete visual occlusion. That's exactly where human video breaks, and exactly what robot policies are missing. We believe cheap, comfortable EMG is the next platform for capturing human interaction: 👓 Egocentric glasses → what the hand sees ⌚️ Wristband → what the hand feels Low-cost, easy to wear, easy to use, scalable. We hope ForceBand sparks more thinking about EMG — and the EMG + Ego form factor — as a new way to capture human data. Vision gave robots the world's appearance — EMG can give them its forces! 🧵 9/n

「🙏 Thank you!!!」 ForceBand is a joint effort across @amazon FAR Team, University of Maryland @UofMaryland , and Johns Hopkins University @JohnsHopkins Huge thanks to our incredible team: @BotaoUMD (Lead), @TX_Leo_Wang , Linna Kuang, Ishaan Ghosh Deepest thanks to the advisors: @JitendraMalikCV , @CFermuller , Tingfan Wu, @maojiayuan , @ruoshi_liu , @HaozhiQ , @YAloimonos 🌐 Website: 📄 Paper: 💻 Code: 🎥 Video: 🧵10/n

Force aware policies is what get us closer to real deployment, nice work!

Thank you, Alfred!

Very nice idea

This is huge

Love this direction! Having inferred force feedback as part of egocentric pre-training would be a massive unlock

This is so cool, it sounds unreal!

Per-finger force from sEMG. 50% lower error than vision baselines. THIS is what robot manipulation learning has needed for YEARS. I need to try this.

Incredible work! This is much better than teleop! I love the idea! Can't wait to see an actual large scale dataset! This + good retargeting + sharpa/wuji-like hand is going to be the future i'm sure!

Great reading. What i'm trying to understand The calibration step still relying on fingertip force sensors. Do you think this is ultimately a data scaling problem (enough diverse sEMG + force pairs), or do you expect some form of calibration to always be necessary across users?

Wrist sensors are cool

Very promising. I think force is the missing variable in much of imitation learning, and a low-cost open-source way to capture it could materially improve robotic manipulation.
