Загрузка видео...

Не удалось загрузить видео

На главную

introducing openego the largest opensource POV manipulation dataset 1107 hours, 120M frames, 290 tasks, 600+ environments task, language actions & 3D hand pose annotations we are collecting egocentric data at scale using AI glasses dm me if your interested in egocentric data

44,008 просмотров • 1 год назад •via X (Twitter)

Комментарии: 36

Фото профиля ahad
ahad1 год назад

currently there is a gap in robotics world models, VLAs, VLMs, and imitation learning require data to generalize we introduce openego to address this the largest collection of open source egocentric footage with fine-grained language and hand joint annotations

Фото профиля ahad
ahad1 год назад

paper: code: website:

Фото профиля ahad
ahad1 год назад

also would like to thank @chris_j_paxton and @svlevine for their blog posts on robotic data that inspired this work and all the authors that open sourced their datasets that make up openego

Фото профиля Bercan
Bercan1 год назад

Wait a second so you didnt collect any new data ?

Фото профиля ahad
ahad1 год назад

yup, two of the dataset included was by our lab @IRVLUTD we are trying to make one place for ego manipulation data similar to open x-embodiment where anyone can contribute

Фото профиля Bercan
Bercan1 год назад

@IRVLUTD Great initiative. I love it, but maybe in your post, make it clear it became only clear when reading the paper.

Фото профиля Alberto Hojel
Alberto Hojel1 год назад

0:10s “right thumb pushes the power button…” is actually index finger VLM captioning ain’t there yet 😓

Фото профиля ahad
ahad1 год назад

Yup, suprisingly a lot of this little errors you can address them with a better prompt.

Фото профиля maxleedev
maxleedev1 год назад

this is sick

Фото профиля ahad
ahad1 год назад

thanks! was inspired by your post 😂😂

Фото профиля Bruno Santos🇵🇹
Bruno Santos🇵🇹1 год назад

@YuXiang_IRVL @pablovelagomez1

Фото профиля Pranav Rapelli
Pranav Rapelli1 год назад

great stuff

Фото профиля Chuck Petras
Chuck Petras1 год назад

@BrianRoemmele

Фото профиля Diann💎🌹
Diann💎🌹1 год назад

@TheShamdoo 😈 Fade hype, but never fade MONSTA.

Фото профиля ζ Pedram ζ
ζ Pedram ζ1 год назад

Very cool, this is the way towards Embodied AI

Фото профиля ahad
ahad1 год назад

Yup, I believe egocentric videos is the key to scaling robot foundation models

Фото профиля towerofshadow
towerofshadow1 год назад

this looks sick ahad

Фото профиля ahad
ahad1 год назад

thanks!

Фото профиля James Han
James Han1 год назад

goat

Фото профиля Avi
Avi1 год назад

so goated

Фото профиля Saïd Aitmbarek
Saïd Aitmbarek1 год назад

based dataset, sick work! feel free to push it to mate

Фото профиля Mario Sonna
Mario Sonna1 год назад

Hi, we need egocentric data, how we can contact you

Фото профиля ahad
ahad1 год назад

sure, if you want access to openego. It’s available on the website And if you want other data just reach out to me either on x or through my email at [email protected]

Фото профиля Boban Jankovic
Boban Jankovic1 год назад

any plans to expand beyond AI glasses for data collection?

Фото профиля ahad
ahad1 год назад

Yeah, we are collecting some data with gopro, and intel realsense cameras hooked up to a headset. Similar to some of the data in the dataset.

Фото профиля Chetan
Chetan1 год назад

cool dataset. have you guys trained anything using it yet? any worries that occlusion causes critical hand tracking issues?

Фото профиля ahad
ahad1 год назад

we have some experiments on trajectory prediction of hand joints. we are planning to do more experiments with it specifically for world models and IL. we have the binary visibility for when it’s occluded. for hand poses we predict we use a confidence threshold on the hand pose.

Фото профиля J
J1 год назад

🔥🔥🔥

Фото профиля Lukas Die Kunst
Lukas Die Kunst1 год назад

This dataset's 120M frames could transform aerodynamics validation. Have you explored applications in sports engineering or UAV simulation?

Фото профиля ahad
ahad1 год назад

this is interesting, I haven't personally thought much on this. how can it be used?

Фото профиля Lukas Die Kunst
Lukas Die Kunst1 год назад

Your dataset could revolutionize aerodynamic cycling gear design - UAE sports tech hubs would jump at this application.

Фото профиля Ran Cheng
Ran Cheng1 год назад

what glasses are you using? project Aria2?

Фото профиля ahad
ahad1 год назад

we built in house glasses

Фото профиля Brooke Gardner
Brooke Gardner1 год назад

Wow, 600+ environments! 🤯 How do AI glasses impact data collection in such diverse settings?

Фото профиля Vishnu
Vishnu1 год назад

Woah this is cool

Фото профиля Riley Ng
Riley Ng10 месяцев назад

congrats man loved how you are collecting egocentric data at scale.

Похожие видео

🔥 JUST IN: Open-source robotics dataset from 100% real-world scenarios! 🤯 Chinese robotics company AGIBOT just released AGIBOT WORLD 2026, an open-source dataset systematically covering key embodied AI research directions. Built entirely from real-world environments: commercial spaces, and homes. Collected using AGIBOT G2 robots in free-form collection mode, providing structured, accurately annotated, high-quality data. Digital twin technology creates 1:1 scale replicas in simulation matching the real environments. Both real-world and simulation data are open-sourced. The AGIBOT G2 platform collects multiple data types simultaneously: RGB(D) cameras, tactile sensors, force sensors, LiDAR, IMU, and full-body joint states. Whole-body control coordinates arms, waist, and hands for complex tasks. First-person teleoperation lets operators control the robot from its perspective. The tasks covered are fine-grained manipulation, ultra-long-horizon tasks, spatial navigation, dual-arm coordination, and multi-agent/human-robot collaboration. The dataset includes error-recovery trajectories with annotations. Most datasets only show successful demonstrations. AGIBOT includes failures and how the robot recovers, teaching models how to handle mistakes. After collection, data is tested through policy training and real-robot deployment to ensure quality. Then processed through industrial quality control with multiple screening and cleaning rounds. Making it open-source accelerates embodied AI research by giving researchers access to high-quality real-world robot data at scale. 🇨🇳 Learn more here: ~~ ♻️ Join the weekly robotics newsletter, and never miss any news →

Lukas Ziegler

40,583 просмотров • 6 месяцев назад

Excited to announce GR00T N1, the world’s first open foundation model for humanoid robots! We are on a mission to democratize Physical AI. The power of general robot brain, in the palm of your hand - with only 2B parameters, N1 learns from the most diverse physical action dataset ever compiled and punches above its weight: - Real humanoid teleoperation data. - Large-scale simulation data: we are open-sourcing 300K+ trajectories! - Neural trajectories: we apply SOTA video generation models to “hallucinate” new synthetic data that features accurate physics in pixels. Using Jensen’s words, “systematically infinite data”! - Latent actions: we develop novel algorithms to extract action tokens from in-the-wild human videos and neural generated videos. GR00T N1 is a single end-to-end neural net, from photons to actions: - Vision-Language Model (System 2) that interprets the physical world through vision and language instructions, enabling robots to reason about their environment and instructions, and plan the right actions. - Diffusion Transformer (System 1) that “renders” smooth and precise motor actions at 120 Hz, executing the latent plan made by System 2. We deploy N1 on GR1 robot, 1X Neo robot, and a large collection of simulation benchmarks. N1 achieves up to +30% boost in diverse manipulation tasks for household and industrial settings. While humanoid robots are the main focus of N1, our model also supports cross-embodiment. We finetune it to work on the $110 HuggingFace LeRobot SO100 robot arm! Open robot brain runs on open hardware. Sounds just right. Let’s solve robotics, together, one token at a time. Links to our Whitepaper, Github repo, HuggingFace model, and open dataset page in the thread: 🧵

Jim Fan

467,596 просмотров • 1 год назад

3D-LLM: Injecting the 3D World into Large Language Models paper page: Large language models (LLMs) and Vision-Language Models (VLMs) have been proven to excel at multiple tasks, such as commonsense reasoning. Powerful as these models can be, they are not grounded in the 3D physical world, which involves richer concepts such as spatial relationships, affordances, physics, layout, and so on. In this work, we propose to inject the 3D world into large language models and introduce a whole new family of 3D-LLMs. Specifically, 3D-LLMs can take 3D point clouds and their features as input and perform a diverse set of 3D-related tasks, including captioning, dense captioning, 3D question answering, task decomposition, 3D grounding, 3D-assisted dialog, navigation, and so on. Using three types of prompting mechanisms that we design, we are able to collect over 300k 3D-language data covering these tasks. To efficiently train 3D-LLMs, we first utilize a 3D feature extractor that obtains 3D features from rendered multi- view images. Then, we use 2D VLMs as our backbones to train our 3D-LLMs. By introducing a 3D localization mechanism, 3D-LLMs can better capture 3D spatial information. Experiments on ScanQA show that our model outperforms state-of-the-art baselines by a large margin (e.g., the BLEU-1 score surpasses state-of-the-art score by 9%). Furthermore, experiments on our held-in datasets for 3D captioning, task composition, and 3D-assisted dialogue show that our model outperforms 2D VLMs. Qualitative examples also show that our model could perform more tasks beyond the scope of existing LLMs and VLMs.

AK

249,798 просмотров • 3 лет назад