Loading video...

Video Failed to Load

Go Home

EmbeddingGemma is our new best-in-class open embedding model designed for on-device AI. 📱 At just 308M parameters, it delivers state-of-the-art performance while being small and efficient enough to run anywhere - even without an internet connection.

584,819 views • 1 year ago •via X (Twitter)

33 Comments

Google DeepMind's profile picture
Google DeepMind1 year ago

🏆 Highest ranking on the MTEB benchmark - the gold standard for text embedding evaluation 🌐 Trained across 100+ languages 🛠️ Ready to go with @huggingface, @llama_index, @langchain and more. Here's how developers can get started with EmbeddingGemma →

JK's profile picture
JK1 year ago

The real breakthrough here is the "without an internet connection" part. On device processing is the key to making AI truly ubiquitous and private. The challenge will be maintaining this level of performance as the model is constrained by ever smaller hardware footprints.

Philip Kiely's profile picture
Philip Kiely1 year ago

Awesome performance and multilingual capabilities in 300M parameters. We have day zero support on Baseten for high-throughput, low-latency deployments:

Prudent AI's profile picture
Prudent AI1 year ago

Google needs better marketing team I guess. @demishassabis why don't you guys make any noise?

Alex Veremeyenko's profile picture
Alex Veremeyenko1 year ago

embedding models getting smaller is a huge win for edge ai. excited to see how developers leverage this offline capability for privacy and efficiency.

JAKARIYA's profile picture
JAKARIYA1 year ago

Love this direction. True intelligence won’t come just from more data or bigger models, but from systems that can exist, adapt, and want to survive in their environment. On-device models feel like the first step toward that.

GraphAI's profile picture
GraphAI1 year ago

Efficiency at this size isn’t just optimization, it’s what makes contextual AI deployable anywhere

Shubham Saboo's profile picture
Shubham Saboo1 year ago

It's amazing. Thank you!

Himanshu Kumar's profile picture
Himanshu Kumar1 year ago

Smaller models, bigger possibilities. Offline accessibility is a game changer for AI's future. This empowers edge computing in exciting new ways.

Tam Nguyen's profile picture
Tam Nguyen1 year ago

Deepmind is cooking

Send it Highor's profile picture
Send it Highor1 year ago

@demishassabis u mad about smol parameters bro? chad move = holding $TROLL 😈

Aryan Agarwal's profile picture
Aryan Agarwal1 year ago

when are u making a better battery life in the pixel by the way please fix that shit 😔😔😔😔😔

Vastkind's profile picture
Vastkind1 year ago

What kinds of apps could flourish once intelligence runs offline as easily as online?

Alexa | Indie hacker's profile picture
Alexa | Indie hacker1 year ago

always good to see your updates, Google 💌 following along with interest

PromptSin's profile picture
PromptSin1 year ago

Wow, that sleek on-device interface is a game-changer for AI accessibility! 🚀

C12s's profile picture
C12s1 year ago

AI that actually fits in your pocket.

Anurag Pant's profile picture
Anurag Pant1 year ago

👍👍👍

Sanchit monga's profile picture
Sanchit monga6 months ago

Love seeing models designed for on-device from the start. 308M params with best-in-class performance is the sweet spot for real-world edge deployment. Smaller, efficient models like this are the future. We're building the runtime layer for deploying them across iOS/Android/Mac/IoT at @RunAnywhereAI.

Mykhailo Sorochuk's profile picture
Mykhailo Sorochuk1 year ago

Edge AI win. Small size, big impact.

OUT_OF EARTH's profile picture
OUT_OF EARTH1 year ago

@grok i am new to ai tools, so can you tell me. What it actually is? What it actually do? What this post is? What this post talking about?

Aasi Tahir Siddique's profile picture
Aasi Tahir Siddique1 year ago

Amazing

Masih Moafi's profile picture
Masih Moafi1 year ago

what do you mean open?! Is it free?

AI PlanetX's profile picture
AI PlanetX1 year ago

Impressive model, compact yet powerful!

Caden Marlow's profile picture
Caden Marlow1 year ago

@grok is there similar available model ? Can i install it on my macbook pro ? Is it already pre-trained so i can discuss technical problem and brainstorm business ideas ? what are the top 3 in device model ?

Alankar Shukla's profile picture
Alankar Shukla1 year ago

Is there any benchmark that represents the performance between gemini embedding model and this gemma model and the openai embedding model

CoralOS's profile picture
CoralOS1 year ago

Impressive work. In Coral, SLM orchestration with small models has already outperformed Microsoft by 34% on the GAIA benchmark, showing how efficiency can beat scale in real tasks. ✅

Dariia Hordiiuk's profile picture
Dariia Hordiiuk1 year ago

On-device power.

saumya 🛠️ truth/acc's profile picture
saumya 🛠️ truth/acc1 year ago

nice

MarginCallMike's profile picture
MarginCallMike1 year ago

$oscr

Pankaj Kumar's profile picture
Pankaj Kumar1 year ago

308M parameters is impressive! Been testing it locally on my old laptop and it's surprisingly fast without sacrificing quality.

Caden Marlow's profile picture
Caden Marlow1 year ago

Did someone tried it on a macbook pro ? Is this the best embedding model ? I was looking for such LLM to work during travel (train, plane ect..) i became so dependant on Cursor, and other AI tools , would be great to have a small LLM for simple task (Brainsorm on code and ideas..)

Timur Klimov's profile picture
Timur Klimov1 year ago

Hey Siri

Jupiter Coder's profile picture
Jupiter Coder1 year ago

On-device AI getting this powerful at just 308M params is mind-blowing 🤯 huge step forward!

Related Videos

Google just proved that bigger isn't always better. Their 308M parameter model is outperforming models 2x its size. Google just released 𝗘𝗺𝗯𝗲𝗱𝗱𝗶𝗻𝗴𝗚𝗲𝗺𝗺𝗮, and it's proving that lightweight embedding models can punch way above their weight class. At just 308M parameters (578MB), it's the new state-of-the-art for models under 500M parameters across MTEB multilingual, English, and code benchmarks. But the really impressive part is that it ranks 8th overall on MTEB(Multilingual, v2) - that's 𝟭𝟳 𝗽𝗹𝗮𝗰𝗲𝘀 above the second-best sub-500M model, and it's delivering performance 𝗰𝗼𝗺𝗽𝗮𝗿𝗮𝗯𝗹𝗲 𝘁𝗼 𝗺𝗼𝗱𝗲𝗹𝘀 𝗻𝗲𝗮𝗿𝗹𝘆 𝗱𝗼𝘂𝗯𝗹𝗲 𝗶𝘁𝘀 𝘀𝗶𝘇𝗲. There are three key parts of their training recipe that sets it apart: 𝟭. 𝗘𝗻𝗰𝗼𝗱𝗲𝗿-𝗗𝗲𝗰𝗼𝗱𝗲𝗿 𝗜𝗻𝗶𝘁𝗶𝗮𝗹𝗶𝘇𝗮𝘁𝗶𝗼𝗻 Instead of starting from a decoder-only Gemma 3 model, they first adapted it to encoder-decoder, then used just the encoder. By basing EmbeddingGemma off an LLM that already has world and language understanding, it gives it a stronger starting point. 𝟮. 𝗧𝗵𝗿𝗲𝗲-𝗟𝗼𝘀𝘀 𝗧𝗿𝗮𝗶𝗻𝗶𝗻𝗴 They combine three different loss functions, instead of just having one: • Contrastive loss (NCE) with in-batch negatives and hardness weighting • Spread-out regularization to ensure embeddings utilize the full space (for quantization and ANN retrieval) • Embedding matching distillation from Gemini Embedding - not just learning from relevance scores, but directly aligning the embedding space with the teacher model 𝟯. 𝗠𝗼𝗱𝗲𝗹 𝗦𝗼𝘂𝗽𝗶𝗻𝗴 Rather than just averaging checkpoints from the same training run, they use optimization techniques to find multiple specialized training mixtures. Each mixture creates an "expert" model in different domains, and averaging all their parameters creates a final model that's actually better than individual models. Extras: • Matryoshka embeddings supporting 768, 512, 256, and 128 dimensions • Quantization-aware training - maintains quality even at int4 precision • 100+ languages from Gemma 3 pretraining • Exceptional performance on low-resource languages (check their XTREME-UP results) Is it the absolute best embedding model? No - Gemini Embedding still leads overall. But that's not really the point. EmbeddingGemma proves you can achieve state-of-the-art performance in a small package that's actually deployable on-device, in low-latency applications, and in resource-constrained environments. This makes good embeddings accessible for use cases that I'm seeing more and more: offline applications, privacy-sensitive deployments, and high-throughput scenarios where inference cost actually matters. Full paper: Shoutout to the EmbeddingGemma team at Google DeepMind for this awesome open source work 💙 and to Daniel Williams for helping me with this video! 🫶

Victoria Slocum

21,610 views • 10 months ago