Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Accepted by #CVPR2023! X-Decoder is the FIRST generalist decoder that supports all segmentation tasks (ins/sem/pano/ref) in OPEN VOCABULARY, both inter- AND intra-image VL tasks, and even helps instruct image inpainting/editing! New demo below and more at

51,930 görüntüleme • 3 yıl önce •via X (Twitter)

6 Yorum

Jianwei Yang profil fotoğrafı
Jianwei Yang3 yıl önce

This project was led by our two wonderful interns @xueyanzou1, @ZiYiDou! With joint mentorship from @zhegan4, @LINJIEFUN, @ChunyuanLi, Xiyang Dai, @HarkiratBehl, Jianfeng Wang, and senior advisory from Violet Peng, Lu Yuan, Lijuan Wang, @yong_jae_lee and @JianfengGao0217!

Akarsh G profil fotoğrafı
Akarsh G3 yıl önce

Used your instruct demo. Still not perfect.

Naoto Usuyama profil fotoğrafı
Naoto Usuyama3 yıl önce

Congrats!

Dan Benyamin (Æ) profil fotoğrafı
Dan Benyamin (Æ)3 yıl önce

Cc @levelsio

Akarsh G profil fotoğrafı
Akarsh G3 yıl önce

How is it different from pix2pix?

Jianwei Yang profil fotoğrafı
Jianwei Yang3 yıl önce

We used our x-decoder as a plug into the original pix2pix to make the edit more grounded.

Benzer Videolar

We benchmarked leading multimodal foundation models (GPT-4o, Claude 3.5 Sonnet, Gemini, Llama, etc.) on standard computer vision tasks—from segmentation to surface normal estimation—using standard datasets like COCO and ImageNet. These models have made remarkable progress; however, it is unclear exactly where they stand in terms of understanding vision in detail. Especially when it comes to tasks beyond question-answering. How well do they understand an object's segments or geometry? Our analyses yield an assessment that is quantitatively and qualitatively detailed and is compatible with evaluations developed in the field of computer vision over the past decades. Observed trends: 🔹 The foundation models consistently underperform task-specific SOTA models across all tasks. However, they are respectable generalists, which is remarkable as they are presumably trained primarily on image-text-based tasks. 🔹 They perform semantic tasks notably better than geometric ones. 🔹 GPT-4o performs the best among non-reasoning models, getting the top position in 4 out of 6 tasks. 🔹 Reasoning models, e.g., o3, show improvements in geometric tasks. 🔹 The 'image generation' models, e.g., GPT-40 Image Generation, which have been natively trained multimodally, exhibit quirks. E.g., hallucinated objects, misalignment between the input and output, etc. 🔹 While the prompting techniques affect performance, better models exhibit less sensitivity to variations in prompts. We control for the variance introduced by the prompting methods in our experiments. 🌐 Detailed analyses, visualizations: ⌨️ code: 🧵 1/n

Amir Zamir

73,244 görüntüleme • 1 yıl önce