Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

One less-known fact about glove-based data collection: it produces higher quality data than teleop on contact-rich tasks. Remote teleop can’t provide good force feedback, but gloves do naturally, making tasks like sock folding, which rely on feel, far easier to capture.

186,027 Aufrufe • vor 8 Monaten •via X (Twitter)

0 Kommentare

Keine Kommentare verfügbar

Kommentare vom Original-Post werden hier angezeigt

Ähnliche Videos

Studies have shown ChatGPT outperforms human annotators for Structured Data by about 25% and costs 30x less. 1 In just 2 months, miners on SN33 running ChatGPT without optimization can’t survive. Today we announce SN33 is now ReadyAI to fully align with our mission 👇 SN33 is building a more performant and significantly cheaper alternative to Scale AI Today structured data is performed primarily by human annotation services like Amazon’s Mechanical Turk and Scale AI It is now more important than ever for every business and individual to make their data AI Ready. However, taking unstructured data and making it Structured Data using today’s tools is extremely costly. SN33 revolutionizes this process, unlocking immense opportunities for commercialization. We lay out the vision for it in this detailed blog post: Validators TODAY can monetize access to this structured data pipeline independently, but we’re streamlining this process, launching a frontend soon that any validator can opt into to provide bandwidth. We've received great feedback from the community, recognizing that what we're building goes far beyond Conversational AI. Building the world's largest annotated conversational dataset (which we've already accomplished) is just one of countless real-world applications for SN33's Structured Data pipeline. We're building a decentralized Scale AI, offering a full suite of Structured Data commodities—from text metadata tagging (available today) to fully customizable queries for company-specific data annotation use cases and image metadata tagging coming soon 👀. Thanks for all the feedback! It has been invaluable so keep bringing it to us! 🙏$TAO Openτensor Foundaτion 1 “ChatGPT Outperforms Crowd-Workers for Text-Annotation Tasks” shows “The zero-shot accuracy of ChatGPT exceeds that of crowd-workers by about 25 percentage points on average [...] Moreover, the per-annotation cost of ChatGPT is less than $0.003—about thirty times cheaper than MTurk”

David Fields

13,639 Aufrufe • vor 2 Jahren

Force-sensing fingers! 🧤 Stanford researchers just released UMI-FT, a handheld data collection platform that puts compact six-axis force/torque sensors on each finger, enabling finger-level wrench measurements alongside RGB, depth, and pose data. Many manipulation tasks require careful force modulation: too little force and the task fails, too much and you cause damage. But commercial force/torque sensors are expensive, bulky, and fragile, which has limited large-scale force-aware policy learning. UMI-FT changes the economics. The platform uses an iPhone for RGB vision, ultrawide RGB, depth, and pose via ARKit, with each finger sensorized using a CoinFT sensor to capture per-finger wrench information during manipulation. This multimodal data trains an adaptive compliance policy that predicts position targets, grasp force, and stiffness for execution on standard compliance controllers. The learned policy runs slowest and generates reference targets, while model-based compliance and force controllers provide delicate 6D compliance control and real-time force modulation. They tested on three contact-rich, force-sensitive tasks: whiteboard wiping (locate eraser, grasp, wipe until clean), skewering zucchini (grasp slice firmly, push onto stick until punctured), and lightbulb insertion (grasp bulb, align bayonet pin with socket slit, insert while overcoming spring force, rotate to light up). The results are clear. Policies without compliance struggle to modulate contact force and trigger safety faults from excessive force. Policies without force sensing fail to grasp unseen objects or resist reaction forces, causing slippage. Here's the project page: ~~ ♻️ Join the weekly robotics newsletter, and never miss any news →

Lukas Ziegler

12,822 Aufrufe • vor 6 Monaten

This is how ALOHA's "teleoperation" system works - a fancy word for "remote control". Training robots will be more and more like playing games in the physical world. A human operates a "joystick++" to perform tasks and collect data, or intervene if there's any safety concern. There's actually a learning curve to master the controller, much like practicing gaming skills. Teleoperation can be done in many different ways. ALOHA is an impressive custom-built system with very low cost. Here're a few alternatives: (1) Motion Capture (MoCap): apply the MoCap systems used for Hollywood movies to capture the fine-grained motions of hand joints. There would be no "embodiment gap" if the robot hand has 5 fingers. For instance, a demonstrator can wear a CyberGlove ( and manipulate the objects. CyberGlove will capture the motion signals & haptic feedback in real-time, which can be re-targeted onto the humanoid. (2) Wearing gloves & markers can be clumsy. An alternative way to do MoCap is through computer vision. DexPilot from NVIDIA enables marker-less and glove-free data collection. The human operator simply uses their bare hands to perform the tasks. 4 Intel RealSense depth cameras and 2 NVIDIA Titan XP GPUs (yeah, 2019 work) translate the pixels to precise motion signals for robot learning. (3) VR Headset: turn the training room into a VR game and "role play" the robot. This has the advantage of scalable remote data collection - annotators from around the world can contribute without coming onsite. VR demonstration technique appeared in research projects like the iGibson home robot simulator, an initiative that I participated in at Stanford: Behind-the-scene video by Litian Liang

Jim Fan

124,588 Aufrufe • vor 2 Jahren

I was really impressed by the UMI gripper (Cheng Chi et al.), but a key limitation is that **force-related data wasn’t captured**: humans feel haptic feedback through the mechanical springs, but the robot couldn’t leverage that info, limiting the data’s value for fine-grained manipulation tasks. Led by my amazing students Yolanda Zhu and Binghao Huang, we designed a **portable visuo-tactile gripper** by integrating our dense, flexible tactile arrays with the UMI gripper to enable large-scale in-the-wild data collection. 🔗 We demonstrate **cross-modal representation learning** and **downstream policy learning** on tasks requiring in-hand state estimation (e.g., test tube reorientation) and fine-grained force sensing (e.g., pipette fluid transfer). Key takeaways: - Our flexible tactile arrays store the rich haptic information humans perceive as dense tactile signals. - Portability and robustness are key for in-the-wild data collection; our portable gripper is compact, lightweight, and durable. - Touch provides precise, robust measurements of in-hand object pose, invariant to lighting and viewpoint. - Cross-modal pretraining on large-scale in-the-wild data significantly improves policy robustness and sample efficiency (as shown many times before — and verified again here!). Also check out our previous investigations of dense, flexible tactile grids for understanding human-robot-environment interactions: - Dense tactile glove (Nature ’19): - 3D-ViTac (CoRL ’24):

Yunzhu Li

13,188 Aufrufe • vor 1 Jahr

is our AI project to make computing feel more human L A N D E R Here are the 4 best demo videos of the magic of DATA in action. DATA is a personalized assistant who knows and remembers every conversation you have with it accross your iPhone, Mac, iPad, Watch, Texts, Emails, and HomePods. You can talk to DATA right in your AirPods or text it just like a person. DATA can read, write, understand, speak any language, and translate between them. It can help with real work and home life tasks like research, writing, scheduling, reminders, and triage. And it's easily customizable so you can have DATA automatically do whatever you want whenever you want with just a few taps and natural language instructions - no code required. DATA can do just about anything you can do on your phone on your behalf automatically including very advanced things Siri can't, like summarizing, analyzing, and drafting replies or writing documents. It can read web pages, texts or emails you show it, or PDFs of any kind. It can do other real world tasks that require complex analysis and common sense too, like: - figure out where the nearest beach is (even when you're in Colorado) and instantly fetch the current surf report up to the current minute. - summarize and drafting replies to entire email chains - plan out entire work projects or multi-day vacations on your calendar - sketch out ideas for you in picture form or drafting Notion pages with charts and graphs. DATA can also use its own judgement to determine when to run an action or not, even if you've scheduled it, allowing you to make VERY complex automations that require many different inputs to make a decision, like for example: - only opening the blinds on your lunch break if it's sunny out and you're working from home. DATA works natively and easily with Apple HomeKit & other shortcuts. DATA can also take initiative and check in with you throughout the day by voice or text and proactively send messages to you and others on your behalf based on your personal and professional goals, current tasks, and calendar. DATA can integrate with many apps on your phone, and is compatible with multiple large AI language models. I've gotten to make a few demo videos that I think really capture how powerful DATA can be for every day life. Here they are all in one tweet. Make sure your sound is on as you watch them. 1. This is the first demo video I ever made from April 19th, 2023. It walks through all the ways you can interact with and use the DATA shortcuts. Everything from saying "Hey Siri" to tapping on custom apps on your home-screen. 2. The second demo video was made May 5 and is an example use case I made of how commands work - commands allow DATA to actually run actions on your phone like taking pictures and sending messages. This demo shows me taking a picture of an email template, and data drafting an email based on that template. It's gotten much better at realizing when it has just run a command and incorporating that information naturally into the conversation now, especially on GPT-4. 3. This third Commands video, May 12 is a walkthrough of ALL the phone functions that commands allow DATA to do: sending texts and emails, making pictures, seeing pictures, reading things, and scheduling events. Since this video we've added auto-replies to texts and emails, summarizing documents, writing documents, health app data retrieval, web surfing, scheduling alarms, making playlists, and more. 4. This last demo I made today, June 15, shows everything DATA does working in concert to generate a crazy detailed morning briefing with background music - including making a unique playlist and giving a detailed analysis of current events complete with Ski & Surf conditions near me other live information from the internet. So now that you've seen everything DATA can do, what's the coolest feature? What features should we add? What would you use DATA for first?

steve

640,114 Aufrufe • vor 3 Jahren

We believe we’re the first robotics company to demonstrate a robot peeling an apple with dual dexterous human-like hands. This breakthrough closes a key gap in robotics, achieving bimanual, contact-rich manipulation and moving far beyond the limits of simple grippers. 🧵↓ Today’s AI models (VLMs) are excellent at perception but struggle with action. Controlling high-degree-of-freedom hands for tasks like this is incredibly complex, and precise finger-level teleoperation is nearly impossible for humans. Our first step was a shared-autonomy system: rather than controlling every finger, the operator triggers pre-learned skills like a “rotate apple or tennis ball” primitive via a keyboard press or pedal. This makes scalable data collection and RL training possible. How does the AI manage this? We created "MoDE-VLA" (Mixture of Dexterous Experts). It fuses vision, language, force, and touch data by using a team of specialist "experts," making control in high-dimensional spaces stable and effective. The combination of these two innovations allows for seamless, contact-rich manipulation. The human provides high-level guidance, and the robot executes the complex in-hand coordination required. This work paves the way for robots that can safely handle delicate tasks in human environments. Want the full technical details? 📄 Read the full research paper: Visit us at NVIDIA GTC Booth #1838, Hall 3 to learn more! #Robotics #AI #DexterousManipulation #VLA #NVIDIAGTC Nancy Villicaña NVIDIA GTC

Sharpa

20,429 Aufrufe • vor 5 Monaten

Mark Zuckerberg: “You can’t 80/20 everything” When Facebook first launched, a user’s profile included things like the dorm they lived in and the courses they were taking. Paul Graham asks Mark if he thinks Facebook would’ve worked without these features, to which Mark replies: “I remember this early debate that Dustin [Moskovitz] and I had where we had to do some manual work for every school that we released Facebook at. To do that, we went through and parsed the course catalogs at the schools to make sure that the data was clean.” Dustin argued that it would be easier to launch new schools if they didn’t parse these catalogs, while Mark thought this would be an unacceptable drop in quality. “We just had this really long debate about what quality meant for us and the community that we wanted to establish and the culture. In retrospect, maybe it wouldn’t have made a huge difference in how things played out. But it definitely set this tone where there’s a lot of clean data on Facebook, you can rely on it, it feels like a college-specific thing—which was valuable early on for setting the culture.” Mark then offers the following advice to the YC Startup School audience: “In the projects you work on, you will have a lot of similar questions. There’s the famous 80/20 rule where you get 80% of the benefit by doing 20% of the work, but you can’t just 80/20 everything. There have to be certain things that you are just the best at and that you go way further than anyone else on to establish this quality bar and have your product be the best thing that’s out there.” Video source: Y Combinator (2012)

Startup Archive

187,836 Aufrufe • vor 4 Monaten

🦔Workers in India are wearing head-mounted cameras for 12 cents an hour to collect training data for humanoid robots. The footage of them doing everyday tasks like cooking, cleaning, sorting, and walking through public spaces gets sold to robotics companies building the models meant to replace those same kinds of jobs in higher-wage countries. The arrangement has been running for roughly two years. Workers do not own the data, do not get residuals, and in many cases are not told what their footage is being used to train. My Take The workers wearing the cameras live in a country where robotics automation will hit decades later, so they are training their own future replacements at a delay that hides the consequence from them personally. The companies buying the data are mostly US and Chinese, building humanoid robots aimed at warehouses, retail, and service jobs in countries paying $15 to $25 an hour rather than 12 cents. Robotics companies need motion data that mimics how humans actually move through real environments, and synthetic data has not been good enough yet. Paying 12 cents an hour in Bengaluru is cheaper than running motion capture studios in Boston, and it works at scale because the worker absorbs the cost of the camera, the discomfort of wearing it, and the long-term loss of any rights to their own movement data. The robotics labor market that eventually emerges from this footage will displace far more wages than the data collection cost to gather. That is the trade investors funding humanoid robotics startups are betting will pay off, and the workers in the videos are the ones paying the tab up front. Hedgie🤗

Hedgie

161,950 Aufrufe • vor 2 Monaten