正在加载视频...

视频加载失败

EmbeddingGemma is our new best-in-class open embedding model designed for on-device AI. 📱 At just 308M parameters, it delivers state-of-the-art performance while being small and efficient enough to run anywhere - even without an internet connection.

584,819 次观看 • 1 年前 •via X (Twitter)

33 条评论

Google DeepMind 的头像
Google DeepMind1 年前

🏆 Highest ranking on the MTEB benchmark - the gold standard for text embedding evaluation 🌐 Trained across 100+ languages 🛠️ Ready to go with @huggingface, @llama_index, @langchain and more. Here's how developers can get started with EmbeddingGemma →

JK 的头像
JK1 年前

The real breakthrough here is the "without an internet connection" part. On device processing is the key to making AI truly ubiquitous and private. The challenge will be maintaining this level of performance as the model is constrained by ever smaller hardware footprints.

Philip Kiely 的头像
Philip Kiely1 年前

Awesome performance and multilingual capabilities in 300M parameters. We have day zero support on Baseten for high-throughput, low-latency deployments:

Prudent AI 的头像
Prudent AI1 年前

Google needs better marketing team I guess. @demishassabis why don't you guys make any noise?

Alex Veremeyenko 的头像
Alex Veremeyenko1 年前

embedding models getting smaller is a huge win for edge ai. excited to see how developers leverage this offline capability for privacy and efficiency.

JAKARIYA 的头像
JAKARIYA1 年前

Love this direction. True intelligence won’t come just from more data or bigger models, but from systems that can exist, adapt, and want to survive in their environment. On-device models feel like the first step toward that.

GraphAI 的头像
GraphAI1 年前

Efficiency at this size isn’t just optimization, it’s what makes contextual AI deployable anywhere

Shubham Saboo 的头像
Shubham Saboo1 年前

It's amazing. Thank you!

Himanshu Kumar 的头像
Himanshu Kumar1 年前

Smaller models, bigger possibilities. Offline accessibility is a game changer for AI's future. This empowers edge computing in exciting new ways.

Tam Nguyen 的头像
Tam Nguyen1 年前

Deepmind is cooking

Send it Highor 的头像
Send it Highor1 年前

@demishassabis u mad about smol parameters bro? chad move = holding $TROLL 😈

Aryan Agarwal 的头像
Aryan Agarwal1 年前

when are u making a better battery life in the pixel by the way please fix that shit 😔😔😔😔😔

Vastkind 的头像
Vastkind1 年前

What kinds of apps could flourish once intelligence runs offline as easily as online?

Alexa | Indie hacker 的头像
Alexa | Indie hacker1 年前

always good to see your updates, Google 💌 following along with interest

PromptSin 的头像
PromptSin1 年前

Wow, that sleek on-device interface is a game-changer for AI accessibility! 🚀

C12s 的头像
C12s1 年前

AI that actually fits in your pocket.

Anurag Pant 的头像
Anurag Pant1 年前

👍👍👍

Sanchit monga 的头像
Sanchit monga6 个月前

Love seeing models designed for on-device from the start. 308M params with best-in-class performance is the sweet spot for real-world edge deployment. Smaller, efficient models like this are the future. We're building the runtime layer for deploying them across iOS/Android/Mac/IoT at @RunAnywhereAI.

Mykhailo Sorochuk 的头像
Mykhailo Sorochuk1 年前

Edge AI win. Small size, big impact.

OUT_OF EARTH 的头像
OUT_OF EARTH1 年前

@grok i am new to ai tools, so can you tell me. What it actually is? What it actually do? What this post is? What this post talking about?

Aasi Tahir Siddique 的头像
Aasi Tahir Siddique1 年前

Amazing

Masih Moafi 的头像
Masih Moafi1 年前

what do you mean open?! Is it free?

AI PlanetX 的头像
AI PlanetX1 年前

Impressive model, compact yet powerful!

Caden Marlow 的头像
Caden Marlow1 年前

@grok is there similar available model ? Can i install it on my macbook pro ? Is it already pre-trained so i can discuss technical problem and brainstorm business ideas ? what are the top 3 in device model ?

Alankar Shukla 的头像
Alankar Shukla1 年前

Is there any benchmark that represents the performance between gemini embedding model and this gemma model and the openai embedding model

CoralOS 的头像
CoralOS1 年前

Impressive work. In Coral, SLM orchestration with small models has already outperformed Microsoft by 34% on the GAIA benchmark, showing how efficiency can beat scale in real tasks. ✅

Dariia Hordiiuk 的头像
Dariia Hordiiuk1 年前

On-device power.

saumya 🛠️ truth/acc 的头像
saumya 🛠️ truth/acc1 年前

nice

MarginCallMike 的头像
MarginCallMike1 年前

$oscr

Pankaj Kumar 的头像
Pankaj Kumar1 年前

308M parameters is impressive! Been testing it locally on my old laptop and it's surprisingly fast without sacrificing quality.

Caden Marlow 的头像
Caden Marlow1 年前

Did someone tried it on a macbook pro ? Is this the best embedding model ? I was looking for such LLM to work during travel (train, plane ect..) i became so dependant on Cursor, and other AI tools , would be great to have a small LLM for simple task (Brainsorm on code and ideas..)

Timur Klimov 的头像
Timur Klimov1 年前

Hey Siri

Jupiter Coder 的头像
Jupiter Coder1 年前

On-device AI getting this powerful at just 308M params is mind-blowing 🤯 huge step forward!

相关视频

Google just proved that bigger isn't always better. Their 308M parameter model is outperforming models 2x its size. Google just released 𝗘𝗺𝗯𝗲𝗱𝗱𝗶𝗻𝗴𝗚𝗲𝗺𝗺𝗮, and it's proving that lightweight embedding models can punch way above their weight class. At just 308M parameters (578MB), it's the new state-of-the-art for models under 500M parameters across MTEB multilingual, English, and code benchmarks. But the really impressive part is that it ranks 8th overall on MTEB(Multilingual, v2) - that's 𝟭𝟳 𝗽𝗹𝗮𝗰𝗲𝘀 above the second-best sub-500M model, and it's delivering performance 𝗰𝗼𝗺𝗽𝗮𝗿𝗮𝗯𝗹𝗲 𝘁𝗼 𝗺𝗼𝗱𝗲𝗹𝘀 𝗻𝗲𝗮𝗿𝗹𝘆 𝗱𝗼𝘂𝗯𝗹𝗲 𝗶𝘁𝘀 𝘀𝗶𝘇𝗲. There are three key parts of their training recipe that sets it apart: 𝟭. 𝗘𝗻𝗰𝗼𝗱𝗲𝗿-𝗗𝗲𝗰𝗼𝗱𝗲𝗿 𝗜𝗻𝗶𝘁𝗶𝗮𝗹𝗶𝘇𝗮𝘁𝗶𝗼𝗻 Instead of starting from a decoder-only Gemma 3 model, they first adapted it to encoder-decoder, then used just the encoder. By basing EmbeddingGemma off an LLM that already has world and language understanding, it gives it a stronger starting point. 𝟮. 𝗧𝗵𝗿𝗲𝗲-𝗟𝗼𝘀𝘀 𝗧𝗿𝗮𝗶𝗻𝗶𝗻𝗴 They combine three different loss functions, instead of just having one: • Contrastive loss (NCE) with in-batch negatives and hardness weighting • Spread-out regularization to ensure embeddings utilize the full space (for quantization and ANN retrieval) • Embedding matching distillation from Gemini Embedding - not just learning from relevance scores, but directly aligning the embedding space with the teacher model 𝟯. 𝗠𝗼𝗱𝗲𝗹 𝗦𝗼𝘂𝗽𝗶𝗻𝗴 Rather than just averaging checkpoints from the same training run, they use optimization techniques to find multiple specialized training mixtures. Each mixture creates an "expert" model in different domains, and averaging all their parameters creates a final model that's actually better than individual models. Extras: • Matryoshka embeddings supporting 768, 512, 256, and 128 dimensions • Quantization-aware training - maintains quality even at int4 precision • 100+ languages from Gemma 3 pretraining • Exceptional performance on low-resource languages (check their XTREME-UP results) Is it the absolute best embedding model? No - Gemini Embedding still leads overall. But that's not really the point. EmbeddingGemma proves you can achieve state-of-the-art performance in a small package that's actually deployable on-device, in low-latency applications, and in resource-constrained environments. This makes good embeddings accessible for use cases that I'm seeing more and more: offline applications, privacy-sensitive deployments, and high-throughput scenarios where inference cost actually matters. Full paper: Shoutout to the EmbeddingGemma team at Google DeepMind for this awesome open source work 💙 and to Daniel Williams for helping me with this video! 🫶

Victoria Slocum

21,610 次观看 • 10 个月前