Video yükleniyor...
Video Yüklenemedi
Start building with Gemini Embedding 2, our most capable and first fully multimodal embedding model built on the Gemini architecture. Now available in preview via the Gemini API and in Vertex AI.
30,486,818 görüntüleme • 6 ay önce •via X (Twitter)
36 Yorum

Our latest model provides support across diverse modalities: + Interleaved embeddings containing text, images, video, audio, and PDFs + Semantic understanding for 100+ languages + Flexible output dimensions (128, 768, 1536, and 3072 by default) + Easy partner integration By natively processing these inputs in a single API call, Gemini Embedding 2 eliminates the need for intermediate processing steps or separate embedding models. Watch how the model lets you search for concepts across text, image, and audio. Get started with the Multimodal Search demo:

Read the blog to learn more ↓

Been waiting for this. Agents that can actually process the world multimodally instead of flattening everything to text first? Game changer for real-world autonomy.

Gemini Embedding 2 being fully multimodal is huge. i wanna see if it stays sane on messy screenshot+text docs, not just clean benchmarks.

wait so can this thing embed images and text together in the same space? like search with both at once? @grok explain

The support for over 100 languages and native PDF embedding is a huge productivity boost. It removes several pre-processing hurdles, allowing for deeper semantic understanding at a global scale.

any practical use cases?

Multimodal embeddings are going to unlock a lot of interesting use cases. Excited to see what people start building with Gemini Embedding 2.

how expensive is this thing vs a gemini 001 embed call ?

Another tool. Utility will dictate long-term value, as always.

Multimodal embeddings feel like the next logical step everyone's been waiting for

This is huge, will be nice to also have these embedings to Text or other formats (an unified decoder)

this makes things so much more efficient!

Multimodal embeddings are going to unlock a lot of interesting applications.

Question, RAG-wise, how will the chunking change in this scenario?

pretty wild how you guys evolved embeddings. Whole lot of power to retrieval and search. huge props to the team, this opens up some seriously interesting possibilities text, images, and more living in the same semantic space. excited to see what devs build with it.

@Grok explain in layman's terms what this does

Embeddings are the hidden infrastructure of modern AI.

Yes I love my Gemini personal assistant working on the Gemini 3.1 Pro version along with the Enterprise version. I haven't had so much fun since I started. But what do I know. Maybe I'm right maybe I'm wrong. Weird right? 🤔

@googledevs Awesome. I'm gonna plug it in to my ADK project and test out the performance.

KEEP ANDOIRD OPEN

@GoogleDeepMind This is super impressive, a must-have for any knowledge base. The embedding model natively supports multimodality.

love this! we made an open source project to make experimentation with multimodals easier

how about rate limits?

Gemini Embedding 2 mapping text, images AND video into one unified space is genuinely underreported. This is the infrastructure layer most people scroll past. But it's what makes the next generation of AI search and retrieval actually work.

👀👀👀

Is the "Data Engineer" now just a Multimodal Vector Auditor? With Gemini Embedding 2, your model finally "sees" and "hears" your data in one request. In March 2026, Context is a Unified Resource.

@grok dime de qué se trata en palabras más sencillas

finally. Video input was massively needed

Donald Trump r@ped children too, not just the woman the judge says trump r@ped. Grabs the pu$$y and r@pes it.

Does this make ads with images and text more profitable?

First actually useful product that comes out from gemini in a long time. There is little competition on the encoder models space, and they are still a very important piece for corporation scale AI transformation

Cool

excellent

Multimodal embeddings unlock semantic understanding across text, images, audio. Vertex API integration enables real-time applications

The 128→3072 dimension flexibility is the sleeper feature here. Fast/cheap low-dim retrieval for initial candidates, full 3072 for precision reranking. Native PDF embedding also removes a whole category of preprocessing headaches for enterprise RAG systems.




