Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

🚀Excited to release PartField—a feedforward model that learns part-based feature fields for 3D shapes! It enables lightning-fast⚡️, robust, open-world hierarchical 3D part seg and unlocks cross-shape applications like co-seg and correspondence! 🔗 1/n

17,292 Aufrufe • vor 1 Jahr •via X (Twitter)

9 Kommentare

Profilbild von Minghua Liu
Minghua Liuvor 1 Jahr

3D part seg remains an open challenge in computer vision. Recent open-world methods distill 2D priors via per-shape optimization but suffer from lengthy runtimes and noisy results. PartField is a feedforward model that delivers lightning-fast speed and robust performance. 2/n

Profilbild von Minghua Liu
Minghua Liuvor 1 Jahr

Instead of relying on text prompts, PartField converts any 3D shape into a feature field that captures the general concept of 3D parts. It then decomposes the shape into hierarchical, multi-granularity parts by applying a clustering algorithm to the field. 3/n

Profilbild von Minghua Liu
Minghua Liuvor 1 Jahr

We train PartField at scale with contrastive learning on both 2D data (distilled masks) and 3D supervision (when available). It showcases strong open-world capabilities across diverse categories, 3D modalities (meshes, Gaussians), and shape styles (artistic, Gen AI, CAD). 4/n

Profilbild von Minghua Liu
Minghua Liuvor 1 Jahr

Instead of recognizing a single part or producing a fixed-granularity clustering, PartField implicitly learns a hierarchy of multi-scale parts and outputs a part tree. Users can interactively choose branches to decompose further based on their desired granularity. 5/n

Profilbild von Minghua Liu
Minghua Liuvor 1 Jahr

Another interesting point is that, while we do not explicitly incorporate any cross-shape supervision, consistency surprisingly emerges in the learned feature space across different shapes. The figure visualizes similarities across the field relative to a selected location. 6/n

Profilbild von Minghua Liu
Minghua Liuvor 1 Jahr

This emergent consistency enables various cross-shape applications, such as shape co-segmentation and correspondence, and demonstrates that PartField learns open-world, general-purpose, hierarchical, and consistent 3D feature fields. 7/n

Profilbild von Minghua Liu
Minghua Liuvor 1 Jahr

Check out our project page, released code, and checkpoints! 🔗 Many thanks to our amazing collaborators: @mikacuy, @DonglaiXiang , @haosu_twitr , @FidlerSanja , @nmwsharp and @JunGao33210520! 8/n

Profilbild von Coinage
Coinagevor 2 Jahren

The first community-owned crypto media outlet has partnered with DAIC, a leading Web3 infrastructure & non-custodial staking provider, to pioneer a new community validator model. Stake with us.

Profilbild von Rawlala
Rawlalavor 1 Jahr

Have u tried with ai genwrated mesh ? 😂

Ähnliche Videos

3D-LLM: Injecting the 3D World into Large Language Models paper page: Large language models (LLMs) and Vision-Language Models (VLMs) have been proven to excel at multiple tasks, such as commonsense reasoning. Powerful as these models can be, they are not grounded in the 3D physical world, which involves richer concepts such as spatial relationships, affordances, physics, layout, and so on. In this work, we propose to inject the 3D world into large language models and introduce a whole new family of 3D-LLMs. Specifically, 3D-LLMs can take 3D point clouds and their features as input and perform a diverse set of 3D-related tasks, including captioning, dense captioning, 3D question answering, task decomposition, 3D grounding, 3D-assisted dialog, navigation, and so on. Using three types of prompting mechanisms that we design, we are able to collect over 300k 3D-language data covering these tasks. To efficiently train 3D-LLMs, we first utilize a 3D feature extractor that obtains 3D features from rendered multi- view images. Then, we use 2D VLMs as our backbones to train our 3D-LLMs. By introducing a 3D localization mechanism, 3D-LLMs can better capture 3D spatial information. Experiments on ScanQA show that our model outperforms state-of-the-art baselines by a large margin (e.g., the BLEU-1 score surpasses state-of-the-art score by 9%). Furthermore, experiments on our held-in datasets for 3D captioning, task composition, and 3D-assisted dialogue show that our model outperforms 2D VLMs. Qualitative examples also show that our model could perform more tasks beyond the scope of existing LLMs and VLMs.

AK

249,798 Aufrufe • vor 3 Jahren