正在加载视频...

视频加载失败

introducing openego the largest opensource POV manipulation dataset 1107 hours, 120M frames, 290 tasks, 600+ environments task, language actions & 3D hand pose annotations we are collecting egocentric data at scale using AI glasses dm me if your interested in egocentric data

44,008 次观看 • 1 年前 •via X (Twitter)

36 条评论

ahad 的头像
ahad1 年前

currently there is a gap in robotics world models, VLAs, VLMs, and imitation learning require data to generalize we introduce openego to address this the largest collection of open source egocentric footage with fine-grained language and hand joint annotations

ahad 的头像
ahad1 年前

paper: code: website:

ahad 的头像
ahad1 年前

also would like to thank @chris_j_paxton and @svlevine for their blog posts on robotic data that inspired this work and all the authors that open sourced their datasets that make up openego

Bercan 的头像
Bercan1 年前

Wait a second so you didnt collect any new data ?

ahad 的头像
ahad1 年前

yup, two of the dataset included was by our lab @IRVLUTD we are trying to make one place for ego manipulation data similar to open x-embodiment where anyone can contribute

Bercan 的头像
Bercan1 年前

@IRVLUTD Great initiative. I love it, but maybe in your post, make it clear it became only clear when reading the paper.

Alberto Hojel 的头像
Alberto Hojel1 年前

0:10s “right thumb pushes the power button…” is actually index finger VLM captioning ain’t there yet 😓

ahad 的头像
ahad1 年前

Yup, suprisingly a lot of this little errors you can address them with a better prompt.

maxleedev 的头像
maxleedev1 年前

this is sick

ahad 的头像
ahad1 年前

thanks! was inspired by your post 😂😂

Bruno Santos🇵🇹 的头像
Bruno Santos🇵🇹1 年前

@YuXiang_IRVL @pablovelagomez1

Pranav Rapelli 的头像
Pranav Rapelli1 年前

great stuff

Chuck Petras 的头像
Chuck Petras1 年前

@BrianRoemmele

Diann💎🌹 的头像
Diann💎🌹1 年前

@TheShamdoo 😈 Fade hype, but never fade MONSTA.

ζ Pedram ζ 的头像
ζ Pedram ζ1 年前

Very cool, this is the way towards Embodied AI

ahad 的头像
ahad1 年前

Yup, I believe egocentric videos is the key to scaling robot foundation models

towerofshadow 的头像
towerofshadow1 年前

this looks sick ahad

ahad 的头像
ahad1 年前

thanks!

James Han 的头像
James Han1 年前

goat

Avi 的头像
Avi1 年前

so goated

Saïd Aitmbarek 的头像
Saïd Aitmbarek1 年前

based dataset, sick work! feel free to push it to mate

Mario Sonna 的头像
Mario Sonna1 年前

Hi, we need egocentric data, how we can contact you

ahad 的头像
ahad1 年前

sure, if you want access to openego. It’s available on the website And if you want other data just reach out to me either on x or through my email at [email protected]

Boban Jankovic 的头像
Boban Jankovic1 年前

any plans to expand beyond AI glasses for data collection?

ahad 的头像
ahad1 年前

Yeah, we are collecting some data with gopro, and intel realsense cameras hooked up to a headset. Similar to some of the data in the dataset.

Chetan 的头像
Chetan1 年前

cool dataset. have you guys trained anything using it yet? any worries that occlusion causes critical hand tracking issues?

ahad 的头像
ahad1 年前

we have some experiments on trajectory prediction of hand joints. we are planning to do more experiments with it specifically for world models and IL. we have the binary visibility for when it’s occluded. for hand poses we predict we use a confidence threshold on the hand pose.

J 的头像
J1 年前

🔥🔥🔥

Lukas Die Kunst 的头像
Lukas Die Kunst1 年前

This dataset's 120M frames could transform aerodynamics validation. Have you explored applications in sports engineering or UAV simulation?

ahad 的头像
ahad1 年前

this is interesting, I haven't personally thought much on this. how can it be used?

Lukas Die Kunst 的头像
Lukas Die Kunst1 年前

Your dataset could revolutionize aerodynamic cycling gear design - UAE sports tech hubs would jump at this application.

Ran Cheng 的头像
Ran Cheng1 年前

what glasses are you using? project Aria2?

ahad 的头像
ahad1 年前

we built in house glasses

Brooke Gardner 的头像
Brooke Gardner1 年前

Wow, 600+ environments! 🤯 How do AI glasses impact data collection in such diverse settings?

Vishnu 的头像
Vishnu1 年前

Woah this is cool

Riley Ng 的头像
Riley Ng10 个月前

congrats man loved how you are collecting egocentric data at scale.

相关视频

🔥 JUST IN: Open-source robotics dataset from 100% real-world scenarios! 🤯 Chinese robotics company AGIBOT just released AGIBOT WORLD 2026, an open-source dataset systematically covering key embodied AI research directions. Built entirely from real-world environments: commercial spaces, and homes. Collected using AGIBOT G2 robots in free-form collection mode, providing structured, accurately annotated, high-quality data. Digital twin technology creates 1:1 scale replicas in simulation matching the real environments. Both real-world and simulation data are open-sourced. The AGIBOT G2 platform collects multiple data types simultaneously: RGB(D) cameras, tactile sensors, force sensors, LiDAR, IMU, and full-body joint states. Whole-body control coordinates arms, waist, and hands for complex tasks. First-person teleoperation lets operators control the robot from its perspective. The tasks covered are fine-grained manipulation, ultra-long-horizon tasks, spatial navigation, dual-arm coordination, and multi-agent/human-robot collaboration. The dataset includes error-recovery trajectories with annotations. Most datasets only show successful demonstrations. AGIBOT includes failures and how the robot recovers, teaching models how to handle mistakes. After collection, data is tested through policy training and real-robot deployment to ensure quality. Then processed through industrial quality control with multiple screening and cleaning rounds. Making it open-source accelerates embodied AI research by giving researchers access to high-quality real-world robot data at scale. 🇨🇳 Learn more here: ~~ ♻️ Join the weekly robotics newsletter, and never miss any news →

Lukas Ziegler

40,583 次观看 • 6 个月前

Excited to announce GR00T N1, the world’s first open foundation model for humanoid robots! We are on a mission to democratize Physical AI. The power of general robot brain, in the palm of your hand - with only 2B parameters, N1 learns from the most diverse physical action dataset ever compiled and punches above its weight: - Real humanoid teleoperation data. - Large-scale simulation data: we are open-sourcing 300K+ trajectories! - Neural trajectories: we apply SOTA video generation models to “hallucinate” new synthetic data that features accurate physics in pixels. Using Jensen’s words, “systematically infinite data”! - Latent actions: we develop novel algorithms to extract action tokens from in-the-wild human videos and neural generated videos. GR00T N1 is a single end-to-end neural net, from photons to actions: - Vision-Language Model (System 2) that interprets the physical world through vision and language instructions, enabling robots to reason about their environment and instructions, and plan the right actions. - Diffusion Transformer (System 1) that “renders” smooth and precise motor actions at 120 Hz, executing the latent plan made by System 2. We deploy N1 on GR1 robot, 1X Neo robot, and a large collection of simulation benchmarks. N1 achieves up to +30% boost in diverse manipulation tasks for household and industrial settings. While humanoid robots are the main focus of N1, our model also supports cross-embodiment. We finetune it to work on the $110 HuggingFace LeRobot SO100 robot arm! Open robot brain runs on open hardware. Sounds just right. Let’s solve robotics, together, one token at a time. Links to our Whitepaper, Github repo, HuggingFace model, and open dataset page in the thread: 🧵

Jim Fan

467,596 次观看 • 1 年前

3D-LLM: Injecting the 3D World into Large Language Models paper page: Large language models (LLMs) and Vision-Language Models (VLMs) have been proven to excel at multiple tasks, such as commonsense reasoning. Powerful as these models can be, they are not grounded in the 3D physical world, which involves richer concepts such as spatial relationships, affordances, physics, layout, and so on. In this work, we propose to inject the 3D world into large language models and introduce a whole new family of 3D-LLMs. Specifically, 3D-LLMs can take 3D point clouds and their features as input and perform a diverse set of 3D-related tasks, including captioning, dense captioning, 3D question answering, task decomposition, 3D grounding, 3D-assisted dialog, navigation, and so on. Using three types of prompting mechanisms that we design, we are able to collect over 300k 3D-language data covering these tasks. To efficiently train 3D-LLMs, we first utilize a 3D feature extractor that obtains 3D features from rendered multi- view images. Then, we use 2D VLMs as our backbones to train our 3D-LLMs. By introducing a 3D localization mechanism, 3D-LLMs can better capture 3D spatial information. Experiments on ScanQA show that our model outperforms state-of-the-art baselines by a large margin (e.g., the BLEU-1 score surpasses state-of-the-art score by 9%). Furthermore, experiments on our held-in datasets for 3D captioning, task composition, and 3D-assisted dialogue show that our model outperforms 2D VLMs. Qualitative examples also show that our model could perform more tasks beyond the scope of existing LLMs and VLMs.

AK

249,798 次观看 • 3 年前