Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Yay, finally! Introducing Vision Banana🍌 from Google DeepMind, our unified model that outperforms SoTA specialist models on various vision tasks! By treating 2D/3D vision tasks as image generation, we unlock a new foundation for CV. Project page: (1/5)

289,558 Aufrufe • vor 4 Monaten •via X (Twitter)

45 Kommentare

Profilbild von Songyou Peng
Songyou Pengvor 4 Monaten

❓How does it work? Take Nano Banana Pro, and instruction-tune it with a mix of: • its original generative data • a small amount of vision task data The key idea: 👉 Represent ALL outputs as RGB images 👉 Control entirely via text prompts One of my favorite example: (2/5)

Profilbild von Songyou Peng
Songyou Pengvor 4 Monaten

What's surprising: Vision Banana keeps its original image generation ability AND achieves state-of-the-art zero-shot performance across tasks. 👉 No task-specific heads. 👉 No special losses. (Yes, the boring table below👇) (3/5)

Profilbild von Songyou Peng
Songyou Pengvor 4 Monaten

The bigger picture: Image generators already master visual understanding during its generative pretraining. We only need a tiny bit of task data to unlock it! 👉 Image generation is the universal interface for vision. (4/5)

Profilbild von Songyou Peng
Songyou Pengvor 4 Monaten

Lastly, what a fun ride co-driving this effort w/ @vgabeur and @ShangbangLong! The AMAZING team🩷 @PaulVoigtlaend1 @Kevin_SSY @jon_barron @NithishKannen @YandongLi8 @bigmoonmandybot @suhasyogin @Yiming20 @huizhong_c @oliver_wang2 @sainingxie @howardzzh @jalayrac @RSoricut (5/5)

Profilbild von Anand Bhattad
Anand Bhattadvor 4 Monaten

@GoogleDeepMind "new" foundation 😉

Profilbild von Songyou Peng
Songyou Pengvor 4 Monaten

@GoogleDeepMind Hi Anand, a really cool work from you (as always)! The key thing in our work is "how" we formulate each task as image generation, without harming the original generation capabilities, and still get SoTA performances across many vision tasks. No specialized loss needed either.

Profilbild von Mike Roberts
Mike Robertsvor 4 Monaten

@anand_bhattad @GoogleDeepMind This couple of sentences is helpful for me as a novice reader! I already made a mental note to ask you about this @songyoupeng and @anand_bhattad, so I’m happy to find this thread ☺️

Profilbild von Anand Bhattad
Anand Bhattadvor 4 Monaten

@songyoupeng @GoogleDeepMind

Profilbild von Khai Loong Aw
Khai Loong Awvor 4 Monaten

Great to see the approach of a unified model for several tasks! (1) What does the term "zero-shot" mean in your post, if it is finetuned on the same task that it is evaluated on — doesn't "zero-shot" usually mean generalizing to unseen tasks or at least categories? Maybe "out-of-distribution generalization" is a better term (though it's hard to tell whether Cityscapes- and COCO-like data are in the pretraining set). Some examples of visual world models from our lab that use self-supervised pretraining and generalize to new tasks without additional training (i.e., truly zero-shot): The key idea is that visual concepts can be extracted zero-shot from a world model by exposing the causal structure of the world through approximate causal inference.

Profilbild von Georgia Gkioxari
Georgia Gkioxarivor 4 Monaten

@GoogleDeepMind Awesome!! Curious to see how well the model works on more complex 2D vision task, such as conversational image grounding

Profilbild von Sean Kirmani
Sean Kirmanivor 4 Monaten

@GoogleDeepMind I’m a fan of the metric depth estimation image! 🏃

Profilbild von Songyou Peng
Songyou Pengvor 4 Monaten

@GoogleDeepMind :))))) 🏃💨

Profilbild von Lior Yariv
Lior Yarivvor 4 Monaten

@GoogleDeepMind So cool!! Have you tried NVS? 👀

Profilbild von Songyou Peng
Songyou Pengvor 4 Monaten

@GoogleDeepMind Sounds like a big request, we probably really should!

Profilbild von Benlin Liu
Benlin Liuvor 4 Monaten

Huge congrats, this is super exciting! Really cool to see this direction pushed so far. We explored a related idea back in 2023 with Unleashing Text-to-Image Diffusion Models for Visual Perception ( where we studied how to leverage pretrained text-to-image diffusion models for downstream perception tasks. Using Stable Diffusion 1.5, that work already showed how far pretrained image generation models could go for perception, including SOTA results on NYUv2 depth estimation, well before follow-up works like Marigold. Excited to see this line of work evolve at GDM 🚀

Profilbild von Stan Szymanowicz
Stan Szymanowiczvor 4 Monaten

@GoogleDeepMind Great stuff @songyoupeng and the entire team!

Profilbild von chrissy
chrissyvor 4 Monaten

@GoogleDeepMind Awesome results!!

Profilbild von Ethan Weber
Ethan Webervor 4 Monaten

@GoogleDeepMind Congrats!! You’ve been cooking. 🔥

Profilbild von Tristan Rhodes
Tristan Rhodesvor 4 Monaten

@GoogleDeepMind Is this going to be a product? An opensource project?

Profilbild von Sherwin Bahmani
Sherwin Bahmanivor 4 Monaten

@GoogleDeepMind Congrats 🥳🚀

Profilbild von Mike Roberts
Mike Robertsvor 4 Monaten

@GoogleDeepMind Yay congratulations! Those surface normal results are extra groovy 🤓🎶

Profilbild von Viv
Vivvor 4 Monaten

@CSProfKGD @GoogleDeepMind awesome work y’all! the table vs SAM3 woah 😮 any example/snippets on how/if you’re post-processing the generated image? ex: can someone extract each instance of an object? any thoughts on extending to video? 👀

Profilbild von Khiem Vuong
Khiem Vuongvor 4 Monaten

@GoogleDeepMind Great work @songyoupeng! Results on in-the-wild examples looks amazing! I’m curious about the evaluation -- did you check for potential data leakage (e.g., whether the base model might have seen any of the evaluation data during pretraining)?

Profilbild von Avais Aziz
Avais Azizvor 4 Monaten

@GoogleDeepMind This is a really elegant way to bridge generation and understanding. The zero-shot results without any task heads are striking.

Profilbild von Ratnesh Madaan
Ratnesh Madaanvor 4 Monaten

@GoogleDeepMind This is super awesome, I had tried a similar prompts sometime back, but specifically to "cubify-ing" objects in any image a few months back. Nano banana can also one shot 3D object detection as well!

Profilbild von Towaki Takikawa / 瀧川永遠希
Towaki Takikawa / 瀧川永遠希vor 4 Monaten

@GoogleDeepMind Super cool!

Profilbild von MaybeRichard
MaybeRichardvor 4 Monaten

@GoogleDeepMind Nice work! Is the code available?

Profilbild von Justin.md
Justin.mdvor 4 Monaten

@GoogleDeepMind Wait can i use this for blender so Gemini can recognize what's wrong with a 3d model and have the precise location to fix it?

Profilbild von Ewan Pedersen
Ewan Pedersenvor 4 Monaten

@GoogleDeepMind this stuff is so fascinating

Profilbild von Rehan Sheikh
Rehan Sheikhvor 4 Monaten

@yongyuanxi @GoogleDeepMind Can this same approach be extended to Veo for video tasks?

Profilbild von Jeya Maria Jose
Jeya Maria Josevor 4 Monaten

@GoogleDeepMind Great work! Congrats @songyoupeng !

Profilbild von JC Rodriguez
JC Rodriguezvor 4 Monaten

@GoogleDeepMind omfg @Swanagan

Profilbild von S
Svor 4 Monaten

@GoogleDeepMind absolutely incredible! how big is the model? and will it be available via an api :)

Profilbild von Yuval Avidani (יובל אבידני)
Yuval Avidani (יובל אבידני)vor 4 Monaten

@ellatechie @GoogleDeepMind How does this help us? What can we do with this capability?

Profilbild von Nikolaos Sarafianos
Nikolaos Sarafianosvor 4 Monaten

@GoogleDeepMind Great work @songyoupeng you've been cooking there!!🧑‍🍳

Profilbild von Songyou Peng
Songyou Pengvor 4 Monaten

@GoogleDeepMind 🧑‍🍳🧑‍🍳🧑‍🍳

Profilbild von Yiming
Yimingvor 4 Monaten

@GoogleDeepMind Excellent work Songyou!

Profilbild von Songyou Peng
Songyou Pengvor 4 Monaten

@GoogleDeepMind Thank YOU Yiming for your continuous support along the journey to make this work possible!!

Profilbild von Abhishek Ramnath
Abhishek Ramnathvor 4 Monaten

@GoogleDeepMind Does it beat metas sam at segmentation?

Profilbild von Yasser Dahou
Yasser Dahouvor 4 Monaten

@GoogleDeepMind amazing work ! also please check our Falcon Perception

Profilbild von Startracker 🔺
Startracker 🔺vor 4 Monaten

@GoogleDeepMind This was the missing piece of the puzzle. Can't wait to give it a spin!

Profilbild von George Birbilis
George Birbilisvor 4 Monaten

@GoogleDeepMind Amazing work

Profilbild von Monter
Montervor 4 Monaten

@GoogleDeepMind Super cool. Already beginning to think of applicable use cases for any of my products 🤔

Profilbild von abi
abivor 4 Monaten

@GoogleDeepMind It's really cool, how people can use it?

Profilbild von Hussain | AI Automation
Hussain | AI Automationvor 1 Monat

@GoogleDeepMind This is a really exciting approach—treating 2D/3D vision tasks as image generation could open up a whole new paradigm for computer vision. 🔥

Ähnliche Videos

3D-LLM: Injecting the 3D World into Large Language Models paper page: Large language models (LLMs) and Vision-Language Models (VLMs) have been proven to excel at multiple tasks, such as commonsense reasoning. Powerful as these models can be, they are not grounded in the 3D physical world, which involves richer concepts such as spatial relationships, affordances, physics, layout, and so on. In this work, we propose to inject the 3D world into large language models and introduce a whole new family of 3D-LLMs. Specifically, 3D-LLMs can take 3D point clouds and their features as input and perform a diverse set of 3D-related tasks, including captioning, dense captioning, 3D question answering, task decomposition, 3D grounding, 3D-assisted dialog, navigation, and so on. Using three types of prompting mechanisms that we design, we are able to collect over 300k 3D-language data covering these tasks. To efficiently train 3D-LLMs, we first utilize a 3D feature extractor that obtains 3D features from rendered multi- view images. Then, we use 2D VLMs as our backbones to train our 3D-LLMs. By introducing a 3D localization mechanism, 3D-LLMs can better capture 3D spatial information. Experiments on ScanQA show that our model outperforms state-of-the-art baselines by a large margin (e.g., the BLEU-1 score surpasses state-of-the-art score by 9%). Furthermore, experiments on our held-in datasets for 3D captioning, task composition, and 3D-assisted dialogue show that our model outperforms 2D VLMs. Qualitative examples also show that our model could perform more tasks beyond the scope of existing LLMs and VLMs.

AK

249,798 Aufrufe • vor 3 Jahren

We benchmarked leading multimodal foundation models (GPT-4o, Claude 3.5 Sonnet, Gemini, Llama, etc.) on standard computer vision tasks—from segmentation to surface normal estimation—using standard datasets like COCO and ImageNet. These models have made remarkable progress; however, it is unclear exactly where they stand in terms of understanding vision in detail. Especially when it comes to tasks beyond question-answering. How well do they understand an object's segments or geometry? Our analyses yield an assessment that is quantitatively and qualitatively detailed and is compatible with evaluations developed in the field of computer vision over the past decades. Observed trends: 🔹 The foundation models consistently underperform task-specific SOTA models across all tasks. However, they are respectable generalists, which is remarkable as they are presumably trained primarily on image-text-based tasks. 🔹 They perform semantic tasks notably better than geometric ones. 🔹 GPT-4o performs the best among non-reasoning models, getting the top position in 4 out of 6 tasks. 🔹 Reasoning models, e.g., o3, show improvements in geometric tasks. 🔹 The 'image generation' models, e.g., GPT-40 Image Generation, which have been natively trained multimodally, exhibit quirks. E.g., hallucinated objects, misalignment between the input and output, etc. 🔹 While the prompting techniques affect performance, better models exhibit less sensitivity to variations in prompts. We control for the variance introduced by the prompting methods in our experiments. 🌐 Detailed analyses, visualizations: ⌨️ code: 🧵 1/n

Amir Zamir

73,244 Aufrufe • vor 1 Jahr