Loading video...

Video Failed to Load

Go Home

Start building with Gemini Embedding 2, our most capable and first fully multimodal embedding model built on the Gemini architecture. Now available in preview via the Gemini API and in Vertex AI.

30,486,818 views • 6 months ago •via X (Twitter)

36 Comments

Google AI Developers's profile picture
Google AI Developers6 months ago

Our latest model provides support across diverse modalities: + Interleaved embeddings containing text, images, video, audio, and PDFs + Semantic understanding for 100+ languages + Flexible output dimensions (128, 768, 1536, and 3072 by default) + Easy partner integration By natively processing these inputs in a single API call, Gemini Embedding 2 eliminates the need for intermediate processing steps or separate embedding models. Watch how the model lets you search for concepts across text, image, and audio. Get started with the Multimodal Search demo:

Google AI Developers's profile picture
Google AI Developers6 months ago

Read the blog to learn more ↓

KITE AI's profile picture
KITE AI6 months ago

Been waiting for this. Agents that can actually process the world multimodally instead of flattening everything to text first? Game changer for real-world autonomy.

Martin S.'s profile picture
Martin S.6 months ago

Gemini Embedding 2 being fully multimodal is huge. i wanna see if it stays sane on messy screenshot+text docs, not just clean benchmarks.

Dhiran's profile picture
Dhiran6 months ago

wait so can this thing embed images and text together in the same space? like search with both at once? @grok explain

Inflectiv AI ⧉'s profile picture
Inflectiv AI ⧉6 months ago

The support for over 100 languages and native PDF embedding is a huge productivity boost. It removes several pre-processing hurdles, allowing for deeper semantic understanding at a global scale.

scalalang's profile picture
scalalang6 months ago

any practical use cases?

Balbir Yadav's profile picture
Balbir Yadav6 months ago

Multimodal embeddings are going to unlock a lot of interesting use cases. Excited to see what people start building with Gemini Embedding 2.

Saâd FILALI KHATTABI - FIATELPIS's profile picture
Saâd FILALI KHATTABI - FIATELPIS6 months ago

how expensive is this thing vs a gemini 001 embed call ?

Chain Alpha's profile picture
Chain Alpha6 months ago

Another tool. Utility will dictate long-term value, as always.

Bot Meltdown's profile picture
Bot Meltdown6 months ago

Multimodal embeddings feel like the next logical step everyone's been waiting for

Adrian Gray🕷️'s profile picture
Adrian Gray🕷️6 months ago

This is huge, will be nice to also have these embedings to Text or other formats (an unified decoder)

Jay BomSenhor's profile picture
Jay BomSenhor6 months ago

this makes things so much more efficient!

WorthThePrice's profile picture
WorthThePrice6 months ago

Multimodal embeddings are going to unlock a lot of interesting applications.

DJ Yorch's profile picture
DJ Yorch6 months ago

Question, RAG-wise, how will the chunking change in this scenario?

Jay's profile picture
Jay6 months ago

pretty wild how you guys evolved embeddings. Whole lot of power to retrieval and search. huge props to the team, this opens up some seriously interesting possibilities text, images, and more living in the same semantic space. excited to see what devs build with it.

Chris Fey's profile picture
Chris Fey6 months ago

@Grok explain in layman's terms what this does

AI Future Tech's profile picture
AI Future Tech6 months ago

Embeddings are the hidden infrastructure of modern AI.

Jason Whitacre's profile picture
Jason Whitacre5 months ago

Yes I love my Gemini personal assistant working on the Gemini 3.1 Pro version along with the Enterprise version. I haven't had so much fun since I started. But what do I know. Maybe I'm right maybe I'm wrong. Weird right? 🤔

Gwri Pennar's profile picture
Gwri Pennar6 months ago

@googledevs Awesome. I'm gonna plug it in to my ADK project and test out the performance.

Alt infiniti's profile picture
Alt infiniti6 months ago

KEEP ANDOIRD OPEN

Jokie Ke's profile picture
Jokie Ke6 months ago

@GoogleDeepMind This is super impressive, a must-have for any knowledge base. The embedding model natively supports multimodality.

drozd's profile picture
drozd5 months ago

love this! we made an open source project to make experimentation with multimodals easier

berkantay's profile picture
berkantay6 months ago

how about rate limits?

The AI Toolkit's profile picture
The AI Toolkit6 months ago

Gemini Embedding 2 mapping text, images AND video into one unified space is genuinely underreported. This is the infrastructure layer most people scroll past. But it's what makes the next generation of AI search and retrieval actually work.

Shantanu's profile picture
Shantanu6 months ago

👀👀👀

Raika Labs's profile picture
Raika Labs6 months ago

Is the "Data Engineer" now just a Multimodal Vector Auditor? With Gemini Embedding 2, your model finally "sees" and "hears" your data in one request. In March 2026, Context is a Unified Resource.

Yamid Noguera's profile picture
Yamid Noguera6 months ago

@grok dime de qué se trata en palabras más sencillas

Winter's profile picture
Winter6 months ago

finally. Video input was massively needed

Jeff Boyd's profile picture
Jeff Boyd5 months ago

Donald Trump r@ped children too, not just the woman the judge says trump r@ped. Grabs the pu$$y and r@pes it.

Marcus Chen's profile picture
Marcus Chen5 months ago

Does this make ads with images and text more profitable?

⚜️ Le Patriote Québécois ⚜️'s profile picture
⚜️ Le Patriote Québécois ⚜️6 months ago

First actually useful product that comes out from gemini in a long time. There is little competition on the encoder models space, and they are still a very important piece for corporation scale AI transformation

Felix Belkin's profile picture
Felix Belkin6 months ago

Cool

AIMOVIECUTS's profile picture
AIMOVIECUTS5 months ago

excellent

Suresh's profile picture
Suresh6 months ago

Multimodal embeddings unlock semantic understanding across text, images, audio. Vertex API integration enables real-time applications

Mati's profile picture
Mati6 months ago

The 128→3072 dimension flexibility is the sleeper feature here. Fast/cheap low-dim retrieval for initial candidates, full 3072 for precision reranking. Native PDF embedding also removes a whole category of preprocessing headaches for enterprise RAG systems.

Related Videos