Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

introducing openego the largest opensource POV manipulation dataset 1107 hours, 120M frames, 290 tasks, 600+ environments task, language actions & 3D hand pose annotations we are collecting egocentric data at scale using AI glasses dm me if your interested in egocentric data

44,008 Aufrufe • vor 1 Jahr •via X (Twitter)

36 Kommentare

Profilbild von ahad
ahadvor 1 Jahr

currently there is a gap in robotics world models, VLAs, VLMs, and imitation learning require data to generalize we introduce openego to address this the largest collection of open source egocentric footage with fine-grained language and hand joint annotations

Profilbild von ahad
ahadvor 1 Jahr

paper: code: website:

Profilbild von ahad
ahadvor 1 Jahr

also would like to thank @chris_j_paxton and @svlevine for their blog posts on robotic data that inspired this work and all the authors that open sourced their datasets that make up openego

Profilbild von Bercan
Bercanvor 1 Jahr

Wait a second so you didnt collect any new data ?

Profilbild von ahad
ahadvor 1 Jahr

yup, two of the dataset included was by our lab @IRVLUTD we are trying to make one place for ego manipulation data similar to open x-embodiment where anyone can contribute

Profilbild von Bercan
Bercanvor 1 Jahr

@IRVLUTD Great initiative. I love it, but maybe in your post, make it clear it became only clear when reading the paper.

Profilbild von Alberto Hojel
Alberto Hojelvor 1 Jahr

0:10s “right thumb pushes the power button…” is actually index finger VLM captioning ain’t there yet 😓

Profilbild von ahad
ahadvor 1 Jahr

Yup, suprisingly a lot of this little errors you can address them with a better prompt.

Profilbild von maxleedev
maxleedevvor 1 Jahr

this is sick

Profilbild von ahad
ahadvor 1 Jahr

thanks! was inspired by your post 😂😂

Profilbild von Bruno Santos🇵🇹
Bruno Santos🇵🇹vor 1 Jahr

@YuXiang_IRVL @pablovelagomez1

Profilbild von Pranav Rapelli
Pranav Rapellivor 1 Jahr

great stuff

Profilbild von Chuck Petras
Chuck Petrasvor 1 Jahr

@BrianRoemmele

Profilbild von Diann💎🌹
Diann💎🌹vor 1 Jahr

@TheShamdoo 😈 Fade hype, but never fade MONSTA.

Profilbild von ζ Pedram ζ
ζ Pedram ζvor 1 Jahr

Very cool, this is the way towards Embodied AI

Profilbild von ahad
ahadvor 1 Jahr

Yup, I believe egocentric videos is the key to scaling robot foundation models

Profilbild von towerofshadow
towerofshadowvor 1 Jahr

this looks sick ahad

Profilbild von ahad
ahadvor 1 Jahr

thanks!

Profilbild von James Han
James Hanvor 1 Jahr

goat

Profilbild von Avi
Avivor 1 Jahr

so goated

Profilbild von Saïd Aitmbarek
Saïd Aitmbarekvor 1 Jahr

based dataset, sick work! feel free to push it to mate

Profilbild von Mario Sonna
Mario Sonnavor 1 Jahr

Hi, we need egocentric data, how we can contact you

Profilbild von ahad
ahadvor 1 Jahr

sure, if you want access to openego. It’s available on the website And if you want other data just reach out to me either on x or through my email at [email protected]

Profilbild von Boban Jankovic
Boban Jankovicvor 1 Jahr

any plans to expand beyond AI glasses for data collection?

Profilbild von ahad
ahadvor 1 Jahr

Yeah, we are collecting some data with gopro, and intel realsense cameras hooked up to a headset. Similar to some of the data in the dataset.

Profilbild von Chetan
Chetanvor 1 Jahr

cool dataset. have you guys trained anything using it yet? any worries that occlusion causes critical hand tracking issues?

Profilbild von ahad
ahadvor 1 Jahr

we have some experiments on trajectory prediction of hand joints. we are planning to do more experiments with it specifically for world models and IL. we have the binary visibility for when it’s occluded. for hand poses we predict we use a confidence threshold on the hand pose.

Profilbild von J
Jvor 1 Jahr

🔥🔥🔥

Profilbild von Lukas Die Kunst
Lukas Die Kunstvor 1 Jahr

This dataset's 120M frames could transform aerodynamics validation. Have you explored applications in sports engineering or UAV simulation?

Profilbild von ahad
ahadvor 1 Jahr

this is interesting, I haven't personally thought much on this. how can it be used?

Profilbild von Lukas Die Kunst
Lukas Die Kunstvor 1 Jahr

Your dataset could revolutionize aerodynamic cycling gear design - UAE sports tech hubs would jump at this application.

Profilbild von Ran Cheng
Ran Chengvor 1 Jahr

what glasses are you using? project Aria2?

Profilbild von ahad
ahadvor 1 Jahr

we built in house glasses

Profilbild von Brooke Gardner
Brooke Gardnervor 1 Jahr

Wow, 600+ environments! 🤯 How do AI glasses impact data collection in such diverse settings?

Profilbild von Vishnu
Vishnuvor 1 Jahr

Woah this is cool

Profilbild von Riley Ng
Riley Ngvor 10 Monaten

congrats man loved how you are collecting egocentric data at scale.

Ähnliche Videos

🔥 JUST IN: Open-source robotics dataset from 100% real-world scenarios! 🤯 Chinese robotics company AGIBOT just released AGIBOT WORLD 2026, an open-source dataset systematically covering key embodied AI research directions. Built entirely from real-world environments: commercial spaces, and homes. Collected using AGIBOT G2 robots in free-form collection mode, providing structured, accurately annotated, high-quality data. Digital twin technology creates 1:1 scale replicas in simulation matching the real environments. Both real-world and simulation data are open-sourced. The AGIBOT G2 platform collects multiple data types simultaneously: RGB(D) cameras, tactile sensors, force sensors, LiDAR, IMU, and full-body joint states. Whole-body control coordinates arms, waist, and hands for complex tasks. First-person teleoperation lets operators control the robot from its perspective. The tasks covered are fine-grained manipulation, ultra-long-horizon tasks, spatial navigation, dual-arm coordination, and multi-agent/human-robot collaboration. The dataset includes error-recovery trajectories with annotations. Most datasets only show successful demonstrations. AGIBOT includes failures and how the robot recovers, teaching models how to handle mistakes. After collection, data is tested through policy training and real-robot deployment to ensure quality. Then processed through industrial quality control with multiple screening and cleaning rounds. Making it open-source accelerates embodied AI research by giving researchers access to high-quality real-world robot data at scale. 🇨🇳 Learn more here: ~~ ♻️ Join the weekly robotics newsletter, and never miss any news →

Lukas Ziegler

40,583 Aufrufe • vor 6 Monaten

Excited to announce GR00T N1, the world’s first open foundation model for humanoid robots! We are on a mission to democratize Physical AI. The power of general robot brain, in the palm of your hand - with only 2B parameters, N1 learns from the most diverse physical action dataset ever compiled and punches above its weight: - Real humanoid teleoperation data. - Large-scale simulation data: we are open-sourcing 300K+ trajectories! - Neural trajectories: we apply SOTA video generation models to “hallucinate” new synthetic data that features accurate physics in pixels. Using Jensen’s words, “systematically infinite data”! - Latent actions: we develop novel algorithms to extract action tokens from in-the-wild human videos and neural generated videos. GR00T N1 is a single end-to-end neural net, from photons to actions: - Vision-Language Model (System 2) that interprets the physical world through vision and language instructions, enabling robots to reason about their environment and instructions, and plan the right actions. - Diffusion Transformer (System 1) that “renders” smooth and precise motor actions at 120 Hz, executing the latent plan made by System 2. We deploy N1 on GR1 robot, 1X Neo robot, and a large collection of simulation benchmarks. N1 achieves up to +30% boost in diverse manipulation tasks for household and industrial settings. While humanoid robots are the main focus of N1, our model also supports cross-embodiment. We finetune it to work on the $110 HuggingFace LeRobot SO100 robot arm! Open robot brain runs on open hardware. Sounds just right. Let’s solve robotics, together, one token at a time. Links to our Whitepaper, Github repo, HuggingFace model, and open dataset page in the thread: 🧵

Jim Fan

467,596 Aufrufe • vor 1 Jahr

3D-LLM: Injecting the 3D World into Large Language Models paper page: Large language models (LLMs) and Vision-Language Models (VLMs) have been proven to excel at multiple tasks, such as commonsense reasoning. Powerful as these models can be, they are not grounded in the 3D physical world, which involves richer concepts such as spatial relationships, affordances, physics, layout, and so on. In this work, we propose to inject the 3D world into large language models and introduce a whole new family of 3D-LLMs. Specifically, 3D-LLMs can take 3D point clouds and their features as input and perform a diverse set of 3D-related tasks, including captioning, dense captioning, 3D question answering, task decomposition, 3D grounding, 3D-assisted dialog, navigation, and so on. Using three types of prompting mechanisms that we design, we are able to collect over 300k 3D-language data covering these tasks. To efficiently train 3D-LLMs, we first utilize a 3D feature extractor that obtains 3D features from rendered multi- view images. Then, we use 2D VLMs as our backbones to train our 3D-LLMs. By introducing a 3D localization mechanism, 3D-LLMs can better capture 3D spatial information. Experiments on ScanQA show that our model outperforms state-of-the-art baselines by a large margin (e.g., the BLEU-1 score surpasses state-of-the-art score by 9%). Furthermore, experiments on our held-in datasets for 3D captioning, task composition, and 3D-assisted dialogue show that our model outperforms 2D VLMs. Qualitative examples also show that our model could perform more tasks beyond the scope of existing LLMs and VLMs.

AK

249,798 Aufrufe • vor 3 Jahren