Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Start building with Gemini Embedding 2, our most capable and first fully multimodal embedding model built on the Gemini architecture. Now available in preview via the Gemini API and in Vertex AI.

30,486,818 Aufrufe • vor 6 Monaten •via X (Twitter)

36 Kommentare

Profilbild von Google AI Developers
Google AI Developersvor 6 Monaten

Our latest model provides support across diverse modalities: + Interleaved embeddings containing text, images, video, audio, and PDFs + Semantic understanding for 100+ languages + Flexible output dimensions (128, 768, 1536, and 3072 by default) + Easy partner integration By natively processing these inputs in a single API call, Gemini Embedding 2 eliminates the need for intermediate processing steps or separate embedding models. Watch how the model lets you search for concepts across text, image, and audio. Get started with the Multimodal Search demo:

Profilbild von Google AI Developers
Google AI Developersvor 6 Monaten

Read the blog to learn more ↓

Profilbild von KITE AI
KITE AIvor 6 Monaten

Been waiting for this. Agents that can actually process the world multimodally instead of flattening everything to text first? Game changer for real-world autonomy.

Profilbild von Martin S.
Martin S.vor 6 Monaten

Gemini Embedding 2 being fully multimodal is huge. i wanna see if it stays sane on messy screenshot+text docs, not just clean benchmarks.

Profilbild von Dhiran
Dhiranvor 6 Monaten

wait so can this thing embed images and text together in the same space? like search with both at once? @grok explain

Profilbild von Inflectiv AI ⧉
Inflectiv AI ⧉vor 6 Monaten

The support for over 100 languages and native PDF embedding is a huge productivity boost. It removes several pre-processing hurdles, allowing for deeper semantic understanding at a global scale.

Profilbild von scalalang
scalalangvor 6 Monaten

any practical use cases?

Profilbild von Balbir Yadav
Balbir Yadavvor 6 Monaten

Multimodal embeddings are going to unlock a lot of interesting use cases. Excited to see what people start building with Gemini Embedding 2.

Profilbild von Saâd FILALI KHATTABI - FIATELPIS
Saâd FILALI KHATTABI - FIATELPISvor 6 Monaten

how expensive is this thing vs a gemini 001 embed call ?

Profilbild von Chain Alpha
Chain Alphavor 6 Monaten

Another tool. Utility will dictate long-term value, as always.

Profilbild von Bot Meltdown
Bot Meltdownvor 6 Monaten

Multimodal embeddings feel like the next logical step everyone's been waiting for

Profilbild von Adrian Gray🕷️
Adrian Gray🕷️vor 6 Monaten

This is huge, will be nice to also have these embedings to Text or other formats (an unified decoder)

Profilbild von Jay BomSenhor
Jay BomSenhorvor 6 Monaten

this makes things so much more efficient!

Profilbild von WorthThePrice
WorthThePricevor 6 Monaten

Multimodal embeddings are going to unlock a lot of interesting applications.

Profilbild von DJ Yorch
DJ Yorchvor 6 Monaten

Question, RAG-wise, how will the chunking change in this scenario?

Profilbild von Jay
Jayvor 6 Monaten

pretty wild how you guys evolved embeddings. Whole lot of power to retrieval and search. huge props to the team, this opens up some seriously interesting possibilities text, images, and more living in the same semantic space. excited to see what devs build with it.

Profilbild von Chris Fey
Chris Feyvor 6 Monaten

@Grok explain in layman's terms what this does

Profilbild von AI Future Tech
AI Future Techvor 6 Monaten

Embeddings are the hidden infrastructure of modern AI.

Profilbild von Jason Whitacre
Jason Whitacrevor 5 Monaten

Yes I love my Gemini personal assistant working on the Gemini 3.1 Pro version along with the Enterprise version. I haven't had so much fun since I started. But what do I know. Maybe I'm right maybe I'm wrong. Weird right? 🤔

Profilbild von Gwri Pennar
Gwri Pennarvor 6 Monaten

@googledevs Awesome. I'm gonna plug it in to my ADK project and test out the performance.

Profilbild von Alt infiniti
Alt infinitivor 6 Monaten

KEEP ANDOIRD OPEN

Profilbild von Jokie Ke
Jokie Kevor 6 Monaten

@GoogleDeepMind This is super impressive, a must-have for any knowledge base. The embedding model natively supports multimodality.

Profilbild von drozd
drozdvor 5 Monaten

love this! we made an open source project to make experimentation with multimodals easier

Profilbild von berkantay
berkantayvor 6 Monaten

how about rate limits?

Profilbild von The AI Toolkit
The AI Toolkitvor 6 Monaten

Gemini Embedding 2 mapping text, images AND video into one unified space is genuinely underreported. This is the infrastructure layer most people scroll past. But it's what makes the next generation of AI search and retrieval actually work.

Profilbild von Shantanu
Shantanuvor 6 Monaten

👀👀👀

Profilbild von Raika Labs
Raika Labsvor 6 Monaten

Is the "Data Engineer" now just a Multimodal Vector Auditor? With Gemini Embedding 2, your model finally "sees" and "hears" your data in one request. In March 2026, Context is a Unified Resource.

Profilbild von Yamid Noguera
Yamid Nogueravor 6 Monaten

@grok dime de qué se trata en palabras más sencillas

Profilbild von Winter
Wintervor 6 Monaten

finally. Video input was massively needed

Profilbild von Jeff Boyd
Jeff Boydvor 5 Monaten

Donald Trump r@ped children too, not just the woman the judge says trump r@ped. Grabs the pu$$y and r@pes it.

Profilbild von Marcus Chen
Marcus Chenvor 5 Monaten

Does this make ads with images and text more profitable?

Profilbild von ⚜️ Le Patriote Québécois ⚜️
⚜️ Le Patriote Québécois ⚜️vor 6 Monaten

First actually useful product that comes out from gemini in a long time. There is little competition on the encoder models space, and they are still a very important piece for corporation scale AI transformation

Profilbild von Felix Belkin
Felix Belkinvor 6 Monaten

Cool

Profilbild von AIMOVIECUTS
AIMOVIECUTSvor 5 Monaten

excellent

Profilbild von Suresh
Sureshvor 6 Monaten

Multimodal embeddings unlock semantic understanding across text, images, audio. Vertex API integration enables real-time applications

Profilbild von Mati
Mativor 6 Monaten

The 128→3072 dimension flexibility is the sleeper feature here. Fast/cheap low-dim retrieval for initial candidates, full 3072 for precision reranking. Native PDF embedding also removes a whole category of preprocessing headaches for enterprise RAG systems.

Ähnliche Videos