正在加载视频...
视频加载失败
It's interesting that the new DeepSeek 4.1 flash is also using very large ngram embeddings ( can be stored on cpu ram ) similar to the upcomming Qwen4 architechture. Here's a fully 3D breakdown of every weight on the model and what it does. Enjoy.
23 条评论

ngram embedding used in Qwen is from a 2026 DeepSeek paper

oooh , which one i missed that one somehow

conditional memory via scalable lookup

How are you so good at Twitter bro?

I live here 😅

What if we find out the engrams can be 99% of a model? RAM gets another 200% more expensive? GPUs at Dollar Tree?

hot

u too

This is so fuggin cool dude

Flash in name, RAM by nature, and now every weight gets a 3D origin story.

This is a good breakdown ! We need to get on a space sometimes soon. Perhaps next week. Anyway listening you on the space, Nisten.

🫡

Nice tool sir 🤝

Awesome!!

Imagine editing these to Your own liking

that's basically how ablitaration of models is done, you check every weight for what it censors and then do precision tuning so you can keep the overall capability loss to a minimum.

full 3d weight breakdown is so cool

ngram as in Engram?

I HAVE BEEN WANTING THIS FOREVER THANK YOU

the full weight map is a great way to make the invisible visible. i keep a small discard list beside experiments; knowing what not to ship is usually the expensive part.

@abidlabs This. Is. Gorgeous.

Wow super interesting!

Great!

