Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

introducing openego the largest opensource POV manipulation dataset 1107 hours, 120M frames, 290 tasks, 600+ environments task, language actions & 3D hand pose annotations we are collecting egocentric data at scale using AI glasses dm me if your interested in egocentric data

44,008 görüntüleme • 1 yıl önce •via X (Twitter)

36 Yorum

ahad profil fotoğrafı
ahad1 yıl önce

currently there is a gap in robotics world models, VLAs, VLMs, and imitation learning require data to generalize we introduce openego to address this the largest collection of open source egocentric footage with fine-grained language and hand joint annotations

ahad profil fotoğrafı
ahad1 yıl önce

paper: code: website:

ahad profil fotoğrafı
ahad1 yıl önce

also would like to thank @chris_j_paxton and @svlevine for their blog posts on robotic data that inspired this work and all the authors that open sourced their datasets that make up openego

Bercan profil fotoğrafı
Bercan1 yıl önce

Wait a second so you didnt collect any new data ?

ahad profil fotoğrafı
ahad1 yıl önce

yup, two of the dataset included was by our lab @IRVLUTD we are trying to make one place for ego manipulation data similar to open x-embodiment where anyone can contribute

Bercan profil fotoğrafı
Bercan1 yıl önce

@IRVLUTD Great initiative. I love it, but maybe in your post, make it clear it became only clear when reading the paper.

Alberto Hojel profil fotoğrafı
Alberto Hojel1 yıl önce

0:10s “right thumb pushes the power button…” is actually index finger VLM captioning ain’t there yet 😓

ahad profil fotoğrafı
ahad1 yıl önce

Yup, suprisingly a lot of this little errors you can address them with a better prompt.

maxleedev profil fotoğrafı
maxleedev1 yıl önce

this is sick

ahad profil fotoğrafı
ahad1 yıl önce

thanks! was inspired by your post 😂😂

Bruno Santos🇵🇹 profil fotoğrafı
Bruno Santos🇵🇹1 yıl önce

@YuXiang_IRVL @pablovelagomez1

Pranav Rapelli profil fotoğrafı
Pranav Rapelli1 yıl önce

great stuff

Chuck Petras profil fotoğrafı
Chuck Petras1 yıl önce

@BrianRoemmele

Diann💎🌹 profil fotoğrafı
Diann💎🌹1 yıl önce

@TheShamdoo 😈 Fade hype, but never fade MONSTA.

ζ Pedram ζ profil fotoğrafı
ζ Pedram ζ1 yıl önce

Very cool, this is the way towards Embodied AI

ahad profil fotoğrafı
ahad1 yıl önce

Yup, I believe egocentric videos is the key to scaling robot foundation models

towerofshadow profil fotoğrafı
towerofshadow1 yıl önce

this looks sick ahad

ahad profil fotoğrafı
ahad1 yıl önce

thanks!

James Han profil fotoğrafı
James Han1 yıl önce

goat

Avi profil fotoğrafı
Avi1 yıl önce

so goated

Saïd Aitmbarek profil fotoğrafı
Saïd Aitmbarek1 yıl önce

based dataset, sick work! feel free to push it to mate

Mario Sonna profil fotoğrafı
Mario Sonna1 yıl önce

Hi, we need egocentric data, how we can contact you

ahad profil fotoğrafı
ahad1 yıl önce

sure, if you want access to openego. It’s available on the website And if you want other data just reach out to me either on x or through my email at [email protected]

Boban Jankovic profil fotoğrafı
Boban Jankovic1 yıl önce

any plans to expand beyond AI glasses for data collection?

ahad profil fotoğrafı
ahad1 yıl önce

Yeah, we are collecting some data with gopro, and intel realsense cameras hooked up to a headset. Similar to some of the data in the dataset.

Chetan profil fotoğrafı
Chetan1 yıl önce

cool dataset. have you guys trained anything using it yet? any worries that occlusion causes critical hand tracking issues?

ahad profil fotoğrafı
ahad1 yıl önce

we have some experiments on trajectory prediction of hand joints. we are planning to do more experiments with it specifically for world models and IL. we have the binary visibility for when it’s occluded. for hand poses we predict we use a confidence threshold on the hand pose.

J profil fotoğrafı
J1 yıl önce

🔥🔥🔥

Lukas Die Kunst profil fotoğrafı
Lukas Die Kunst1 yıl önce

This dataset's 120M frames could transform aerodynamics validation. Have you explored applications in sports engineering or UAV simulation?

ahad profil fotoğrafı
ahad1 yıl önce

this is interesting, I haven't personally thought much on this. how can it be used?

Lukas Die Kunst profil fotoğrafı
Lukas Die Kunst1 yıl önce

Your dataset could revolutionize aerodynamic cycling gear design - UAE sports tech hubs would jump at this application.

Ran Cheng profil fotoğrafı
Ran Cheng1 yıl önce

what glasses are you using? project Aria2?

ahad profil fotoğrafı
ahad1 yıl önce

we built in house glasses

Brooke Gardner profil fotoğrafı
Brooke Gardner1 yıl önce

Wow, 600+ environments! 🤯 How do AI glasses impact data collection in such diverse settings?

Vishnu profil fotoğrafı
Vishnu1 yıl önce

Woah this is cool

Riley Ng profil fotoğrafı
Riley Ng10 ay önce

congrats man loved how you are collecting egocentric data at scale.

Benzer Videolar

🔥 JUST IN: Open-source robotics dataset from 100% real-world scenarios! 🤯 Chinese robotics company AGIBOT just released AGIBOT WORLD 2026, an open-source dataset systematically covering key embodied AI research directions. Built entirely from real-world environments: commercial spaces, and homes. Collected using AGIBOT G2 robots in free-form collection mode, providing structured, accurately annotated, high-quality data. Digital twin technology creates 1:1 scale replicas in simulation matching the real environments. Both real-world and simulation data are open-sourced. The AGIBOT G2 platform collects multiple data types simultaneously: RGB(D) cameras, tactile sensors, force sensors, LiDAR, IMU, and full-body joint states. Whole-body control coordinates arms, waist, and hands for complex tasks. First-person teleoperation lets operators control the robot from its perspective. The tasks covered are fine-grained manipulation, ultra-long-horizon tasks, spatial navigation, dual-arm coordination, and multi-agent/human-robot collaboration. The dataset includes error-recovery trajectories with annotations. Most datasets only show successful demonstrations. AGIBOT includes failures and how the robot recovers, teaching models how to handle mistakes. After collection, data is tested through policy training and real-robot deployment to ensure quality. Then processed through industrial quality control with multiple screening and cleaning rounds. Making it open-source accelerates embodied AI research by giving researchers access to high-quality real-world robot data at scale. 🇨🇳 Learn more here: ~~ ♻️ Join the weekly robotics newsletter, and never miss any news →

Lukas Ziegler

40,583 görüntüleme • 6 ay önce

Excited to announce GR00T N1, the world’s first open foundation model for humanoid robots! We are on a mission to democratize Physical AI. The power of general robot brain, in the palm of your hand - with only 2B parameters, N1 learns from the most diverse physical action dataset ever compiled and punches above its weight: - Real humanoid teleoperation data. - Large-scale simulation data: we are open-sourcing 300K+ trajectories! - Neural trajectories: we apply SOTA video generation models to “hallucinate” new synthetic data that features accurate physics in pixels. Using Jensen’s words, “systematically infinite data”! - Latent actions: we develop novel algorithms to extract action tokens from in-the-wild human videos and neural generated videos. GR00T N1 is a single end-to-end neural net, from photons to actions: - Vision-Language Model (System 2) that interprets the physical world through vision and language instructions, enabling robots to reason about their environment and instructions, and plan the right actions. - Diffusion Transformer (System 1) that “renders” smooth and precise motor actions at 120 Hz, executing the latent plan made by System 2. We deploy N1 on GR1 robot, 1X Neo robot, and a large collection of simulation benchmarks. N1 achieves up to +30% boost in diverse manipulation tasks for household and industrial settings. While humanoid robots are the main focus of N1, our model also supports cross-embodiment. We finetune it to work on the $110 HuggingFace LeRobot SO100 robot arm! Open robot brain runs on open hardware. Sounds just right. Let’s solve robotics, together, one token at a time. Links to our Whitepaper, Github repo, HuggingFace model, and open dataset page in the thread: 🧵

Jim Fan

467,596 görüntüleme • 1 yıl önce

3D-LLM: Injecting the 3D World into Large Language Models paper page: Large language models (LLMs) and Vision-Language Models (VLMs) have been proven to excel at multiple tasks, such as commonsense reasoning. Powerful as these models can be, they are not grounded in the 3D physical world, which involves richer concepts such as spatial relationships, affordances, physics, layout, and so on. In this work, we propose to inject the 3D world into large language models and introduce a whole new family of 3D-LLMs. Specifically, 3D-LLMs can take 3D point clouds and their features as input and perform a diverse set of 3D-related tasks, including captioning, dense captioning, 3D question answering, task decomposition, 3D grounding, 3D-assisted dialog, navigation, and so on. Using three types of prompting mechanisms that we design, we are able to collect over 300k 3D-language data covering these tasks. To efficiently train 3D-LLMs, we first utilize a 3D feature extractor that obtains 3D features from rendered multi- view images. Then, we use 2D VLMs as our backbones to train our 3D-LLMs. By introducing a 3D localization mechanism, 3D-LLMs can better capture 3D spatial information. Experiments on ScanQA show that our model outperforms state-of-the-art baselines by a large margin (e.g., the BLEU-1 score surpasses state-of-the-art score by 9%). Furthermore, experiments on our held-in datasets for 3D captioning, task composition, and 3D-assisted dialogue show that our model outperforms 2D VLMs. Qualitative examples also show that our model could perform more tasks beyond the scope of existing LLMs and VLMs.

AK

249,798 görüntüleme • 3 yıl önce