Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Google just released PaliGemma 2 Mix: new versatile instruction vision language models 🔥 > Three new models: 3B, 10B, 28B with res 224, 448 💙 > Can do vision language tasks with open-ended prompts, understand documents, and segment or detect anything 🤯

83,712 Aufrufe • vor 1 Jahr •via X (Twitter)

11 Kommentare

Profilbild von merve
mervevor 1 Jahr

All models and the demo are here 🤝 Read our blog to learn more 📚

Profilbild von merve
mervevor 1 Jahr

@onuralpszr this might interest you!

Profilbild von Quantify Funds
Quantify Fundsvor 1 Jahr

Watch this space for the reveal 👀 BIG NEWS: Our family of ETFs is expanding with 10 new single stock offerings. Read the pre-effective prospectus @ #STKd #QuantifyFunds #QuantifyChaos #Nasdaq #ETFs

Profilbild von ondevice
ondevicevor 1 Jahr

love how, usually when we come across something from google, we usually see demos rather than just academic benchmarks thanks team! love this @DynamicWebPaige @OfficialLoganK @AarushSelvan

Profilbild von lambda
lambdavor 1 Jahr

openAI should take notes. Even google is open sourcing

Profilbild von merve
mervevor 1 Jahr

google has been open sourcing since early times and for a long while actually

Profilbild von nonc3ai
nonc3aivor 1 Jahr

Great news, 3B should be small enough to run on a phone, right?

Profilbild von merve
mervevor 1 Jahr

if you quantize, yes ☺️

Profilbild von Arpit Sharma
Arpit Sharmavor 1 Jahr

AI that sees, understands, and segments. Impressive!

Profilbild von Matthew Rogers
Matthew Rogersvor 1 Jahr

I want to scan my mail, lets go!

Profilbild von Furkan Gözükara
Furkan Gözükaravor 1 Jahr

Why res is still to low?

Ähnliche Videos

We benchmarked leading multimodal foundation models (GPT-4o, Claude 3.5 Sonnet, Gemini, Llama, etc.) on standard computer vision tasks—from segmentation to surface normal estimation—using standard datasets like COCO and ImageNet. These models have made remarkable progress; however, it is unclear exactly where they stand in terms of understanding vision in detail. Especially when it comes to tasks beyond question-answering. How well do they understand an object's segments or geometry? Our analyses yield an assessment that is quantitatively and qualitatively detailed and is compatible with evaluations developed in the field of computer vision over the past decades. Observed trends: 🔹 The foundation models consistently underperform task-specific SOTA models across all tasks. However, they are respectable generalists, which is remarkable as they are presumably trained primarily on image-text-based tasks. 🔹 They perform semantic tasks notably better than geometric ones. 🔹 GPT-4o performs the best among non-reasoning models, getting the top position in 4 out of 6 tasks. 🔹 Reasoning models, e.g., o3, show improvements in geometric tasks. 🔹 The 'image generation' models, e.g., GPT-40 Image Generation, which have been natively trained multimodally, exhibit quirks. E.g., hallucinated objects, misalignment between the input and output, etc. 🔹 While the prompting techniques affect performance, better models exhibit less sensitivity to variations in prompts. We control for the variance introduced by the prompting methods in our experiments. 🌐 Detailed analyses, visualizations: ⌨️ code: 🧵 1/n

Amir Zamir

73,244 Aufrufe • vor 1 Jahr