Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Start building with Gemini Embedding 2, our most capable and first fully multimodal embedding model built on the Gemini architecture. Now available in preview via the Gemini API and in Vertex AI.

30,486,818 görüntüleme • 6 ay önce •via X (Twitter)

36 Yorum

Google AI Developers profil fotoğrafı
Google AI Developers6 ay önce

Our latest model provides support across diverse modalities: + Interleaved embeddings containing text, images, video, audio, and PDFs + Semantic understanding for 100+ languages + Flexible output dimensions (128, 768, 1536, and 3072 by default) + Easy partner integration By natively processing these inputs in a single API call, Gemini Embedding 2 eliminates the need for intermediate processing steps or separate embedding models. Watch how the model lets you search for concepts across text, image, and audio. Get started with the Multimodal Search demo:

Google AI Developers profil fotoğrafı
Google AI Developers6 ay önce

Read the blog to learn more ↓

KITE AI profil fotoğrafı
KITE AI6 ay önce

Been waiting for this. Agents that can actually process the world multimodally instead of flattening everything to text first? Game changer for real-world autonomy.

Martin S. profil fotoğrafı
Martin S.6 ay önce

Gemini Embedding 2 being fully multimodal is huge. i wanna see if it stays sane on messy screenshot+text docs, not just clean benchmarks.

Dhiran profil fotoğrafı
Dhiran6 ay önce

wait so can this thing embed images and text together in the same space? like search with both at once? @grok explain

Inflectiv AI ⧉ profil fotoğrafı
Inflectiv AI ⧉6 ay önce

The support for over 100 languages and native PDF embedding is a huge productivity boost. It removes several pre-processing hurdles, allowing for deeper semantic understanding at a global scale.

scalalang profil fotoğrafı
scalalang6 ay önce

any practical use cases?

Balbir Yadav profil fotoğrafı
Balbir Yadav6 ay önce

Multimodal embeddings are going to unlock a lot of interesting use cases. Excited to see what people start building with Gemini Embedding 2.

Saâd FILALI KHATTABI - FIATELPIS profil fotoğrafı
Saâd FILALI KHATTABI - FIATELPIS6 ay önce

how expensive is this thing vs a gemini 001 embed call ?

Chain Alpha profil fotoğrafı
Chain Alpha6 ay önce

Another tool. Utility will dictate long-term value, as always.

Bot Meltdown profil fotoğrafı
Bot Meltdown6 ay önce

Multimodal embeddings feel like the next logical step everyone's been waiting for

Adrian Gray🕷️ profil fotoğrafı
Adrian Gray🕷️6 ay önce

This is huge, will be nice to also have these embedings to Text or other formats (an unified decoder)

Jay BomSenhor profil fotoğrafı
Jay BomSenhor6 ay önce

this makes things so much more efficient!

WorthThePrice profil fotoğrafı
WorthThePrice6 ay önce

Multimodal embeddings are going to unlock a lot of interesting applications.

DJ Yorch profil fotoğrafı
DJ Yorch6 ay önce

Question, RAG-wise, how will the chunking change in this scenario?

Jay profil fotoğrafı
Jay6 ay önce

pretty wild how you guys evolved embeddings. Whole lot of power to retrieval and search. huge props to the team, this opens up some seriously interesting possibilities text, images, and more living in the same semantic space. excited to see what devs build with it.

Chris Fey profil fotoğrafı
Chris Fey6 ay önce

@Grok explain in layman's terms what this does

AI Future Tech profil fotoğrafı
AI Future Tech6 ay önce

Embeddings are the hidden infrastructure of modern AI.

Jason Whitacre profil fotoğrafı
Jason Whitacre5 ay önce

Yes I love my Gemini personal assistant working on the Gemini 3.1 Pro version along with the Enterprise version. I haven't had so much fun since I started. But what do I know. Maybe I'm right maybe I'm wrong. Weird right? 🤔

Gwri Pennar profil fotoğrafı
Gwri Pennar6 ay önce

@googledevs Awesome. I'm gonna plug it in to my ADK project and test out the performance.

Alt infiniti profil fotoğrafı
Alt infiniti6 ay önce

KEEP ANDOIRD OPEN

Jokie Ke profil fotoğrafı
Jokie Ke6 ay önce

@GoogleDeepMind This is super impressive, a must-have for any knowledge base. The embedding model natively supports multimodality.

drozd profil fotoğrafı
drozd5 ay önce

love this! we made an open source project to make experimentation with multimodals easier

berkantay profil fotoğrafı
berkantay6 ay önce

how about rate limits?

The AI Toolkit profil fotoğrafı
The AI Toolkit6 ay önce

Gemini Embedding 2 mapping text, images AND video into one unified space is genuinely underreported. This is the infrastructure layer most people scroll past. But it's what makes the next generation of AI search and retrieval actually work.

Shantanu profil fotoğrafı
Shantanu6 ay önce

👀👀👀

Raika Labs profil fotoğrafı
Raika Labs6 ay önce

Is the "Data Engineer" now just a Multimodal Vector Auditor? With Gemini Embedding 2, your model finally "sees" and "hears" your data in one request. In March 2026, Context is a Unified Resource.

Yamid Noguera profil fotoğrafı
Yamid Noguera6 ay önce

@grok dime de qué se trata en palabras más sencillas

Winter profil fotoğrafı
Winter6 ay önce

finally. Video input was massively needed

Jeff Boyd profil fotoğrafı
Jeff Boyd5 ay önce

Donald Trump r@ped children too, not just the woman the judge says trump r@ped. Grabs the pu$$y and r@pes it.

Marcus Chen profil fotoğrafı
Marcus Chen5 ay önce

Does this make ads with images and text more profitable?

⚜️ Le Patriote Québécois ⚜️ profil fotoğrafı
⚜️ Le Patriote Québécois ⚜️6 ay önce

First actually useful product that comes out from gemini in a long time. There is little competition on the encoder models space, and they are still a very important piece for corporation scale AI transformation

Felix Belkin profil fotoğrafı
Felix Belkin6 ay önce

Cool

AIMOVIECUTS profil fotoğrafı
AIMOVIECUTS5 ay önce

excellent

Suresh profil fotoğrafı
Suresh6 ay önce

Multimodal embeddings unlock semantic understanding across text, images, audio. Vertex API integration enables real-time applications

Mati profil fotoğrafı
Mati6 ay önce

The 128→3072 dimension flexibility is the sleeper feature here. Fast/cheap low-dim retrieval for initial candidates, full 3072 for precision reranking. Native PDF embedding also removes a whole category of preprocessing headaches for enterprise RAG systems.

Benzer Videolar