Video wird geladen...
Video konnte nicht geladen werden
Yay, finally! Introducing Vision Banana🍌 from Google DeepMind, our unified model that outperforms SoTA specialist models on various vision tasks! By treating 2D/3D vision tasks as image generation, we unlock a new foundation for CV. Project page: (1/5)
289,558 Aufrufe • vor 4 Monaten •via X (Twitter)
45 Kommentare

❓How does it work? Take Nano Banana Pro, and instruction-tune it with a mix of: • its original generative data • a small amount of vision task data The key idea: 👉 Represent ALL outputs as RGB images 👉 Control entirely via text prompts One of my favorite example: (2/5)

What's surprising: Vision Banana keeps its original image generation ability AND achieves state-of-the-art zero-shot performance across tasks. 👉 No task-specific heads. 👉 No special losses. (Yes, the boring table below👇) (3/5)

The bigger picture: Image generators already master visual understanding during its generative pretraining. We only need a tiny bit of task data to unlock it! 👉 Image generation is the universal interface for vision. (4/5)

Lastly, what a fun ride co-driving this effort w/ @vgabeur and @ShangbangLong! The AMAZING team🩷 @PaulVoigtlaend1 @Kevin_SSY @jon_barron @NithishKannen @YandongLi8 @bigmoonmandybot @suhasyogin @Yiming20 @huizhong_c @oliver_wang2 @sainingxie @howardzzh @jalayrac @RSoricut (5/5)

@GoogleDeepMind "new" foundation 😉

@GoogleDeepMind Hi Anand, a really cool work from you (as always)! The key thing in our work is "how" we formulate each task as image generation, without harming the original generation capabilities, and still get SoTA performances across many vision tasks. No specialized loss needed either.

@anand_bhattad @GoogleDeepMind This couple of sentences is helpful for me as a novice reader! I already made a mental note to ask you about this @songyoupeng and @anand_bhattad, so I’m happy to find this thread ☺️

@songyoupeng @GoogleDeepMind

Great to see the approach of a unified model for several tasks! (1) What does the term "zero-shot" mean in your post, if it is finetuned on the same task that it is evaluated on — doesn't "zero-shot" usually mean generalizing to unseen tasks or at least categories? Maybe "out-of-distribution generalization" is a better term (though it's hard to tell whether Cityscapes- and COCO-like data are in the pretraining set). Some examples of visual world models from our lab that use self-supervised pretraining and generalize to new tasks without additional training (i.e., truly zero-shot): The key idea is that visual concepts can be extracted zero-shot from a world model by exposing the causal structure of the world through approximate causal inference.

@GoogleDeepMind Awesome!! Curious to see how well the model works on more complex 2D vision task, such as conversational image grounding

@GoogleDeepMind I’m a fan of the metric depth estimation image! 🏃

@GoogleDeepMind :))))) 🏃💨

@GoogleDeepMind So cool!! Have you tried NVS? 👀

@GoogleDeepMind Sounds like a big request, we probably really should!

Huge congrats, this is super exciting! Really cool to see this direction pushed so far. We explored a related idea back in 2023 with Unleashing Text-to-Image Diffusion Models for Visual Perception ( where we studied how to leverage pretrained text-to-image diffusion models for downstream perception tasks. Using Stable Diffusion 1.5, that work already showed how far pretrained image generation models could go for perception, including SOTA results on NYUv2 depth estimation, well before follow-up works like Marigold. Excited to see this line of work evolve at GDM 🚀

@GoogleDeepMind Great stuff @songyoupeng and the entire team!

@GoogleDeepMind Awesome results!!

@GoogleDeepMind Congrats!! You’ve been cooking. 🔥

@GoogleDeepMind Is this going to be a product? An opensource project?

@GoogleDeepMind Congrats 🥳🚀

@GoogleDeepMind Yay congratulations! Those surface normal results are extra groovy 🤓🎶

@CSProfKGD @GoogleDeepMind awesome work y’all! the table vs SAM3 woah 😮 any example/snippets on how/if you’re post-processing the generated image? ex: can someone extract each instance of an object? any thoughts on extending to video? 👀

@GoogleDeepMind Great work @songyoupeng! Results on in-the-wild examples looks amazing! I’m curious about the evaluation -- did you check for potential data leakage (e.g., whether the base model might have seen any of the evaluation data during pretraining)?

@GoogleDeepMind This is a really elegant way to bridge generation and understanding. The zero-shot results without any task heads are striking.

@GoogleDeepMind This is super awesome, I had tried a similar prompts sometime back, but specifically to "cubify-ing" objects in any image a few months back. Nano banana can also one shot 3D object detection as well!

@GoogleDeepMind Super cool!

@GoogleDeepMind Nice work! Is the code available?

@GoogleDeepMind Wait can i use this for blender so Gemini can recognize what's wrong with a 3d model and have the precise location to fix it?

@GoogleDeepMind this stuff is so fascinating

@yongyuanxi @GoogleDeepMind Can this same approach be extended to Veo for video tasks?

@GoogleDeepMind Great work! Congrats @songyoupeng !

@GoogleDeepMind omfg @Swanagan

@GoogleDeepMind absolutely incredible! how big is the model? and will it be available via an api :)

@ellatechie @GoogleDeepMind How does this help us? What can we do with this capability?

@GoogleDeepMind Great work @songyoupeng you've been cooking there!!🧑🍳

@GoogleDeepMind 🧑🍳🧑🍳🧑🍳

@GoogleDeepMind Excellent work Songyou!

@GoogleDeepMind Thank YOU Yiming for your continuous support along the journey to make this work possible!!

@GoogleDeepMind Does it beat metas sam at segmentation?

@GoogleDeepMind amazing work ! also please check our Falcon Perception

@GoogleDeepMind This was the missing piece of the puzzle. Can't wait to give it a spin!

@GoogleDeepMind Amazing work

@GoogleDeepMind Super cool. Already beginning to think of applicable use cases for any of my products 🤔

@GoogleDeepMind It's really cool, how people can use it?

@GoogleDeepMind This is a really exciting approach—treating 2D/3D vision tasks as image generation could open up a whole new paradigm for computer vision. 🔥
