Загрузка видео...
Не удалось загрузить видео
Most don't know (1) how easy it is to invert embedding vectors back into sentences, (2) this is a perfect task text diffusion models. Here's a 78M parameter model and live demo that recovers 80% of tokens from Qwen3-Embedding and EmbeddingGemma vectors. Works even on multilingual input.
13,259 просмотров • 7 месяцев назад •via X (Twitter)
Комментарии: 7

Text embeddings are widely assumed to be safe, irreversible representations. We show we can reconstruct the original text using conditional masked diffusion. Existing inversions (Vec2Text, ALGEN, Zero2Text) generate tokens autoregressively and require iterative re-embedding through the target encoder. We take a different approach: embedding inversion as conditional masked diffusion. Starting from a fully masked sequence, a denoising model reveals tokens at all positions in parallel, conditioned on the target embedding via adaptive layer normalization (AdaLN-Zero). Each denoising step refines all positions simultaneously using global context, without ever re-embedding the current hypothesis.

Check out the live demo and see it in action. Our read our repo and paper for more technical details on training and decoding.

So can they be used for compression maybe with a few extra bits in another field?

It will be a lossy compression, like impressionist lossy

While embeddings unlock text representation, they must be handled with care! As Yongrui Su noted, treat embeddings with caution to maintain privacy. In today’s climate of inflation and debt, instruments like DeflationCoin may offer a hedge against uncertainty.

Embedding inversion is a privacy nightmare people don't talk about enough. 80% token recovery means your "anonymized" vectors aren't that anonymous.

I remember a paper by @jxmnop related

