Loading video...

Video Failed to Load

Go Home

introducing openego the largest opensource POV manipulation dataset 1107 hours, 120M frames, 290 tasks, 600+ environments task, language actions & 3D hand pose annotations we are collecting egocentric data at scale using AI glasses dm me if your interested in egocentric data

44,008 views • 1 year ago •via X (Twitter)

36 Comments

ahad's profile picture
ahad1 year ago

currently there is a gap in robotics world models, VLAs, VLMs, and imitation learning require data to generalize we introduce openego to address this the largest collection of open source egocentric footage with fine-grained language and hand joint annotations

ahad's profile picture
ahad1 year ago

paper: code: website:

ahad's profile picture
ahad1 year ago

also would like to thank @chris_j_paxton and @svlevine for their blog posts on robotic data that inspired this work and all the authors that open sourced their datasets that make up openego

Bercan's profile picture
Bercan1 year ago

Wait a second so you didnt collect any new data ?

ahad's profile picture
ahad1 year ago

yup, two of the dataset included was by our lab @IRVLUTD we are trying to make one place for ego manipulation data similar to open x-embodiment where anyone can contribute

Bercan's profile picture
Bercan1 year ago

@IRVLUTD Great initiative. I love it, but maybe in your post, make it clear it became only clear when reading the paper.

Alberto Hojel's profile picture
Alberto Hojel1 year ago

0:10s “right thumb pushes the power button…” is actually index finger VLM captioning ain’t there yet 😓

ahad's profile picture
ahad1 year ago

Yup, suprisingly a lot of this little errors you can address them with a better prompt.

maxleedev's profile picture
maxleedev1 year ago

this is sick

ahad's profile picture
ahad1 year ago

thanks! was inspired by your post 😂😂

Bruno Santos🇵🇹's profile picture
Bruno Santos🇵🇹1 year ago

@YuXiang_IRVL @pablovelagomez1

Pranav Rapelli's profile picture
Pranav Rapelli1 year ago

great stuff

Chuck Petras's profile picture
Chuck Petras1 year ago

@BrianRoemmele

Diann💎🌹's profile picture
Diann💎🌹1 year ago

@TheShamdoo 😈 Fade hype, but never fade MONSTA.

ζ Pedram ζ's profile picture
ζ Pedram ζ1 year ago

Very cool, this is the way towards Embodied AI

ahad's profile picture
ahad1 year ago

Yup, I believe egocentric videos is the key to scaling robot foundation models

towerofshadow's profile picture
towerofshadow1 year ago

this looks sick ahad

ahad's profile picture
ahad1 year ago

thanks!

James Han's profile picture
James Han1 year ago

goat

Avi's profile picture
Avi1 year ago

so goated

Saïd Aitmbarek's profile picture
Saïd Aitmbarek1 year ago

based dataset, sick work! feel free to push it to mate

Mario Sonna's profile picture
Mario Sonna1 year ago

Hi, we need egocentric data, how we can contact you

ahad's profile picture
ahad1 year ago

sure, if you want access to openego. It’s available on the website And if you want other data just reach out to me either on x or through my email at [email protected]

Boban Jankovic's profile picture
Boban Jankovic1 year ago

any plans to expand beyond AI glasses for data collection?

ahad's profile picture
ahad1 year ago

Yeah, we are collecting some data with gopro, and intel realsense cameras hooked up to a headset. Similar to some of the data in the dataset.

Chetan's profile picture
Chetan1 year ago

cool dataset. have you guys trained anything using it yet? any worries that occlusion causes critical hand tracking issues?

ahad's profile picture
ahad1 year ago

we have some experiments on trajectory prediction of hand joints. we are planning to do more experiments with it specifically for world models and IL. we have the binary visibility for when it’s occluded. for hand poses we predict we use a confidence threshold on the hand pose.

J's profile picture
J1 year ago

🔥🔥🔥

Lukas Die Kunst's profile picture
Lukas Die Kunst1 year ago

This dataset's 120M frames could transform aerodynamics validation. Have you explored applications in sports engineering or UAV simulation?

ahad's profile picture
ahad1 year ago

this is interesting, I haven't personally thought much on this. how can it be used?

Lukas Die Kunst's profile picture
Lukas Die Kunst1 year ago

Your dataset could revolutionize aerodynamic cycling gear design - UAE sports tech hubs would jump at this application.

Ran Cheng's profile picture
Ran Cheng1 year ago

what glasses are you using? project Aria2?

ahad's profile picture
ahad1 year ago

we built in house glasses

Brooke Gardner's profile picture
Brooke Gardner1 year ago

Wow, 600+ environments! 🤯 How do AI glasses impact data collection in such diverse settings?

Vishnu's profile picture
Vishnu1 year ago

Woah this is cool

Riley Ng's profile picture
Riley Ng10 months ago

congrats man loved how you are collecting egocentric data at scale.

Related Videos

🔥 JUST IN: Open-source robotics dataset from 100% real-world scenarios! 🤯 Chinese robotics company AGIBOT just released AGIBOT WORLD 2026, an open-source dataset systematically covering key embodied AI research directions. Built entirely from real-world environments: commercial spaces, and homes. Collected using AGIBOT G2 robots in free-form collection mode, providing structured, accurately annotated, high-quality data. Digital twin technology creates 1:1 scale replicas in simulation matching the real environments. Both real-world and simulation data are open-sourced. The AGIBOT G2 platform collects multiple data types simultaneously: RGB(D) cameras, tactile sensors, force sensors, LiDAR, IMU, and full-body joint states. Whole-body control coordinates arms, waist, and hands for complex tasks. First-person teleoperation lets operators control the robot from its perspective. The tasks covered are fine-grained manipulation, ultra-long-horizon tasks, spatial navigation, dual-arm coordination, and multi-agent/human-robot collaboration. The dataset includes error-recovery trajectories with annotations. Most datasets only show successful demonstrations. AGIBOT includes failures and how the robot recovers, teaching models how to handle mistakes. After collection, data is tested through policy training and real-robot deployment to ensure quality. Then processed through industrial quality control with multiple screening and cleaning rounds. Making it open-source accelerates embodied AI research by giving researchers access to high-quality real-world robot data at scale. 🇨🇳 Learn more here: ~~ ♻️ Join the weekly robotics newsletter, and never miss any news →

Lukas Ziegler

40,583 views • 6 months ago

Excited to announce GR00T N1, the world’s first open foundation model for humanoid robots! We are on a mission to democratize Physical AI. The power of general robot brain, in the palm of your hand - with only 2B parameters, N1 learns from the most diverse physical action dataset ever compiled and punches above its weight: - Real humanoid teleoperation data. - Large-scale simulation data: we are open-sourcing 300K+ trajectories! - Neural trajectories: we apply SOTA video generation models to “hallucinate” new synthetic data that features accurate physics in pixels. Using Jensen’s words, “systematically infinite data”! - Latent actions: we develop novel algorithms to extract action tokens from in-the-wild human videos and neural generated videos. GR00T N1 is a single end-to-end neural net, from photons to actions: - Vision-Language Model (System 2) that interprets the physical world through vision and language instructions, enabling robots to reason about their environment and instructions, and plan the right actions. - Diffusion Transformer (System 1) that “renders” smooth and precise motor actions at 120 Hz, executing the latent plan made by System 2. We deploy N1 on GR1 robot, 1X Neo robot, and a large collection of simulation benchmarks. N1 achieves up to +30% boost in diverse manipulation tasks for household and industrial settings. While humanoid robots are the main focus of N1, our model also supports cross-embodiment. We finetune it to work on the $110 HuggingFace LeRobot SO100 robot arm! Open robot brain runs on open hardware. Sounds just right. Let’s solve robotics, together, one token at a time. Links to our Whitepaper, Github repo, HuggingFace model, and open dataset page in the thread: 🧵

Jim Fan

467,596 views • 1 year ago

3D-LLM: Injecting the 3D World into Large Language Models paper page: Large language models (LLMs) and Vision-Language Models (VLMs) have been proven to excel at multiple tasks, such as commonsense reasoning. Powerful as these models can be, they are not grounded in the 3D physical world, which involves richer concepts such as spatial relationships, affordances, physics, layout, and so on. In this work, we propose to inject the 3D world into large language models and introduce a whole new family of 3D-LLMs. Specifically, 3D-LLMs can take 3D point clouds and their features as input and perform a diverse set of 3D-related tasks, including captioning, dense captioning, 3D question answering, task decomposition, 3D grounding, 3D-assisted dialog, navigation, and so on. Using three types of prompting mechanisms that we design, we are able to collect over 300k 3D-language data covering these tasks. To efficiently train 3D-LLMs, we first utilize a 3D feature extractor that obtains 3D features from rendered multi- view images. Then, we use 2D VLMs as our backbones to train our 3D-LLMs. By introducing a 3D localization mechanism, 3D-LLMs can better capture 3D spatial information. Experiments on ScanQA show that our model outperforms state-of-the-art baselines by a large margin (e.g., the BLEU-1 score surpasses state-of-the-art score by 9%). Furthermore, experiments on our held-in datasets for 3D captioning, task composition, and 3D-assisted dialogue show that our model outperforms 2D VLMs. Qualitative examples also show that our model could perform more tasks beyond the scope of existing LLMs and VLMs.

AK

249,798 views • 3 years ago