Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Introducing Meta Locate 3D: a model for accurate object localization in 3D environments. Learn how Meta Locate 3D can help robots accurately understand their surroundings and interact more naturally with humans. You can download the model and dataset, read our research paper, and even try a demo!

81,406 Aufrufe • vor 1 Jahr •via X (Twitter)

11 Kommentare

Profilbild von Arunachalam B
Arunachalam Bvor 1 Jahr

YES! This is exactly what we need! 🙌 Accurate 3D object localization will revolutionize robotics and human-robot interaction! Going to download the model and try it now!

Profilbild von Rainmaker
Rainmakervor 2 Jahren

Which Machine Learning model delivers stronger trading results? Check out this free Substack post where I compare several powerful models that beat the market and show yearly returns of over 20%.

Profilbild von Reji Modiyil
Reji Modiyilvor 1 Jahr

@AIatMeta, exciting advancements like meta locate 3d transform how we connect machines and humans.

Profilbild von -T3ch-
-T3ch-vor 1 Jahr

is this built on meta perception?

Profilbild von Casey Ash
Casey Ashvor 1 Jahr

Nice model for local context; it makes me wonder how it could enable more user-friendly applications.

Profilbild von Melina Flow
Melina Flowvor 1 Jahr

Curious how this model adapts across diverse environments. Will download and evaluate your insights.

Profilbild von Faruk Guney
Faruk Guneyvor 1 Jahr

This could very well be a great addition to SLAM toolbox in robotics if proven effective.

Profilbild von Alejandro Toro
Alejandro Torovor 1 Jahr

Pretty awesome. I have to check it out

Profilbild von TrainLabsAI
TrainLabsAIvor 1 Jahr

THE TRAIN also helps, but it helps humans

Profilbild von o-mega.ai
o-mega.aivor 1 Jahr

Meta Locate 3D revolutionizes 3D object localization using natural language queries. The model's 3D-JEPA algorithm processes 3D point clouds with 2D foundation models like CLIP and DINO, operating on sensor streams for real-time deployment. Backed by a 130,000+ annotation dataset, it's poised to transform robot-human interactions. The open-source codebase on GitHub opens doors for developers to push AR and robotics capabilities further, paving the way for AI systems that navigate and interpret physical spaces based on human instructions. This marks a significant leap towards more intuitive and capable AI-driven environmental understanding.

Profilbild von DeBarra Shaw
DeBarra Shawvor 1 Jahr

Nice.

Ähnliche Videos

3D-LLM: Injecting the 3D World into Large Language Models paper page: Large language models (LLMs) and Vision-Language Models (VLMs) have been proven to excel at multiple tasks, such as commonsense reasoning. Powerful as these models can be, they are not grounded in the 3D physical world, which involves richer concepts such as spatial relationships, affordances, physics, layout, and so on. In this work, we propose to inject the 3D world into large language models and introduce a whole new family of 3D-LLMs. Specifically, 3D-LLMs can take 3D point clouds and their features as input and perform a diverse set of 3D-related tasks, including captioning, dense captioning, 3D question answering, task decomposition, 3D grounding, 3D-assisted dialog, navigation, and so on. Using three types of prompting mechanisms that we design, we are able to collect over 300k 3D-language data covering these tasks. To efficiently train 3D-LLMs, we first utilize a 3D feature extractor that obtains 3D features from rendered multi- view images. Then, we use 2D VLMs as our backbones to train our 3D-LLMs. By introducing a 3D localization mechanism, 3D-LLMs can better capture 3D spatial information. Experiments on ScanQA show that our model outperforms state-of-the-art baselines by a large margin (e.g., the BLEU-1 score surpasses state-of-the-art score by 9%). Furthermore, experiments on our held-in datasets for 3D captioning, task composition, and 3D-assisted dialogue show that our model outperforms 2D VLMs. Qualitative examples also show that our model could perform more tasks beyond the scope of existing LLMs and VLMs.

AK

249,708 Aufrufe • vor 3 Jahren