正在加载视频...

视频加载失败

Start building with Gemini Embedding 2, our most capable and first fully multimodal embedding model built on the Gemini architecture. Now available in preview via the Gemini API and in Vertex AI.

30,486,818 次观看 • 6 个月前 •via X (Twitter)

36 条评论

Google AI Developers 的头像
Google AI Developers6 个月前

Our latest model provides support across diverse modalities: + Interleaved embeddings containing text, images, video, audio, and PDFs + Semantic understanding for 100+ languages + Flexible output dimensions (128, 768, 1536, and 3072 by default) + Easy partner integration By natively processing these inputs in a single API call, Gemini Embedding 2 eliminates the need for intermediate processing steps or separate embedding models. Watch how the model lets you search for concepts across text, image, and audio. Get started with the Multimodal Search demo:

Google AI Developers 的头像
Google AI Developers6 个月前

Read the blog to learn more ↓

KITE AI 的头像
KITE AI6 个月前

Been waiting for this. Agents that can actually process the world multimodally instead of flattening everything to text first? Game changer for real-world autonomy.

Martin S. 的头像
Martin S.6 个月前

Gemini Embedding 2 being fully multimodal is huge. i wanna see if it stays sane on messy screenshot+text docs, not just clean benchmarks.

Dhiran 的头像
Dhiran6 个月前

wait so can this thing embed images and text together in the same space? like search with both at once? @grok explain

Inflectiv AI ⧉ 的头像
Inflectiv AI ⧉6 个月前

The support for over 100 languages and native PDF embedding is a huge productivity boost. It removes several pre-processing hurdles, allowing for deeper semantic understanding at a global scale.

scalalang 的头像
scalalang6 个月前

any practical use cases?

Balbir Yadav 的头像
Balbir Yadav6 个月前

Multimodal embeddings are going to unlock a lot of interesting use cases. Excited to see what people start building with Gemini Embedding 2.

Saâd FILALI KHATTABI - FIATELPIS 的头像
Saâd FILALI KHATTABI - FIATELPIS6 个月前

how expensive is this thing vs a gemini 001 embed call ?

Chain Alpha 的头像
Chain Alpha6 个月前

Another tool. Utility will dictate long-term value, as always.

Bot Meltdown 的头像
Bot Meltdown6 个月前

Multimodal embeddings feel like the next logical step everyone's been waiting for

Adrian Gray🕷️ 的头像
Adrian Gray🕷️6 个月前

This is huge, will be nice to also have these embedings to Text or other formats (an unified decoder)

Jay BomSenhor 的头像
Jay BomSenhor6 个月前

this makes things so much more efficient!

WorthThePrice 的头像
WorthThePrice6 个月前

Multimodal embeddings are going to unlock a lot of interesting applications.

DJ Yorch 的头像
DJ Yorch6 个月前

Question, RAG-wise, how will the chunking change in this scenario?

Jay 的头像
Jay6 个月前

pretty wild how you guys evolved embeddings. Whole lot of power to retrieval and search. huge props to the team, this opens up some seriously interesting possibilities text, images, and more living in the same semantic space. excited to see what devs build with it.

Chris Fey 的头像
Chris Fey6 个月前

@Grok explain in layman's terms what this does

AI Future Tech 的头像
AI Future Tech6 个月前

Embeddings are the hidden infrastructure of modern AI.

Jason Whitacre 的头像
Jason Whitacre5 个月前

Yes I love my Gemini personal assistant working on the Gemini 3.1 Pro version along with the Enterprise version. I haven't had so much fun since I started. But what do I know. Maybe I'm right maybe I'm wrong. Weird right? 🤔

Gwri Pennar 的头像
Gwri Pennar6 个月前

@googledevs Awesome. I'm gonna plug it in to my ADK project and test out the performance.

Alt infiniti 的头像
Alt infiniti6 个月前

KEEP ANDOIRD OPEN

Jokie Ke 的头像
Jokie Ke6 个月前

@GoogleDeepMind This is super impressive, a must-have for any knowledge base. The embedding model natively supports multimodality.

drozd 的头像
drozd5 个月前

love this! we made an open source project to make experimentation with multimodals easier

berkantay 的头像
berkantay6 个月前

how about rate limits?

The AI Toolkit 的头像
The AI Toolkit6 个月前

Gemini Embedding 2 mapping text, images AND video into one unified space is genuinely underreported. This is the infrastructure layer most people scroll past. But it's what makes the next generation of AI search and retrieval actually work.

Shantanu 的头像
Shantanu6 个月前

👀👀👀

Raika Labs 的头像
Raika Labs6 个月前

Is the "Data Engineer" now just a Multimodal Vector Auditor? With Gemini Embedding 2, your model finally "sees" and "hears" your data in one request. In March 2026, Context is a Unified Resource.

Yamid Noguera 的头像
Yamid Noguera6 个月前

@grok dime de qué se trata en palabras más sencillas

Winter 的头像
Winter6 个月前

finally. Video input was massively needed

Jeff Boyd 的头像
Jeff Boyd5 个月前

Donald Trump r@ped children too, not just the woman the judge says trump r@ped. Grabs the pu$$y and r@pes it.

Marcus Chen 的头像
Marcus Chen5 个月前

Does this make ads with images and text more profitable?

⚜️ Le Patriote Québécois ⚜️ 的头像
⚜️ Le Patriote Québécois ⚜️6 个月前

First actually useful product that comes out from gemini in a long time. There is little competition on the encoder models space, and they are still a very important piece for corporation scale AI transformation

Felix Belkin 的头像
Felix Belkin6 个月前

Cool

AIMOVIECUTS 的头像
AIMOVIECUTS5 个月前

excellent

Suresh 的头像
Suresh6 个月前

Multimodal embeddings unlock semantic understanding across text, images, audio. Vertex API integration enables real-time applications

Mati 的头像
Mati6 个月前

The 128→3072 dimension flexibility is the sleeper feature here. Fast/cheap low-dim retrieval for initial candidates, full 3072 for precision reranking. Native PDF embedding also removes a whole category of preprocessing headaches for enterprise RAG systems.

相关视频