Загрузка видео...

Не удалось загрузить видео

На главную

Sparse attention (MoBA/NSA) trains faster & beats full attention in key tasks. But we’ve had no idea how they truly work…until now. 🔍 We reverse-engineered them to uncover: - Novel attention patterns - Hidden "attention sinks" - Better performance - And more A 🧵… ~1/8~

59,480 просмотров • 1 год назад •via X (Twitter)

Комментарии: 11

Фото профиля Tilde
Tilde1 год назад

📖 Read the full post here: ~2/8~ Sparse attention exploits inherent sparsity in model attention patterns to dramatically accelerate sequence mixing. Natively trainable approaches, such as Kimi’s MoBA and Deepseek’s NSA, expand the pareto frontier by matching and even outcompeting base attention on expressivity respectively.

Фото профиля Tilde
Tilde1 год назад

~3/8~ We trained dozens of sparse attention models and poked around in their brains. Sparse attention models boost superior long-context generalization capability out of box, even with 80% sparsity in attention scores.

Фото профиля Tilde
Tilde1 год назад

~4/8~ We visualized the first-ever long-context attention maps for sparse attention, revealing fascinating patterns and attention circuits.

Фото профиля Tilde
Tilde1 год назад

~5/8~ We identified the novel mechanisms of attention sinking in sparse attention models. We found MoBA models periodically re-inject attention sinks and that these sinks can be easily identified using value norm and attention score heuristics.

Фото профиля Tilde
Tilde1 год назад

~6/8~ We investigated the key geometry for different attention models, and also found that KV-sharing through GQA strongly influences the resulting query/key manifolds. This offered a clue into how sparse attention could be beating base attention → by removing cross-head interference introduced by GQA!

Фото профиля Tilde
Tilde1 год назад

~7/8~ We analyzed the gating distributions for NSA models and found we can ablate many branches without compromising model performance! Our principled ablations enabled massive gains in throughput without losses in performance.

Фото профиля Tilde
Tilde1 год назад

~8/8~ We release our NSA kernel for experimentation and research here: At Tilde, we believe interpretability is the path towards building better models. If that sounds cool, reach out!

Фото профиля Tilde
Tilde1 год назад

Read the full post here:

Фото профиля Mobile Scanner
Mobile Scanner1 год назад

Scan any documents, convert images into text, PDF files, etc. 👍

Фото профиля Nate Chen
Nate Chen1 год назад

just reached out to tilde's email, when can I hear back and join the hiring process?

Фото профиля Himanshu Kumar
Himanshu Kumar1 год назад

Surprisingly, "attention sinks" might be the key to efficient learning, not a flaw.

Похожие видео

New short course: Attention in Transformers: Concepts and Code in PyTorch. Last week we released a course on how LLM transformers work. This week, go deeper and learn about the technical ideas behind the attention mechanism, and see how to code it in PyTorch. This course is built with Joshua Starmer, Founder and CEO of StatQuest. The attention mechanism was a breakthrough that led to transformers, the architecture powering large language models like ChatGPT. Transformers, introduced in the 2017 paper: "Attention is All You Need" by Viswani and others, took off because of its highly scalable design. In this course, you’ll learn how the attention mechanism, a key element of transformer-based LLMs, works and implement it in PyTorch. You'll develop deep intuition about building reliable, functional, and scalable AI applications. What you will do: - Understand the evolution of the attention mechanism, a key breakthrough that led to transformers. - Learn the relationships between word embeddings, positional embeddings, and attention. - Learn about the Query, Key, and Value matrices, and how to produce and use them in attention. - Walk through the math required to calculate self-attention and masked self-attention to learn why and how they work. - Understand the difference between self-attention and masked self-attention and how one is used in the encoder to build context-aware embeddings and the other is used in the decoder for generative outputs. - Learn the details of the encoder-decoder architecture, cross-attention, and multi-head attention and how they are all incorporated into a transformer. - Use PyTorch to code a class that implements self-attention, masked self-attention, and multi-head attention. There're lots of exciting technical details in this course. Please sign up here:

Andrew Ng

132,357 просмотров • 1 год назад

I hear so often from the Dommes I work with that they struggle with people online fetichizing them and simply seeing them for how sexy and beautiful they are. They project their fantasies and their desires onto you. That stops immediately once you move the attention from you to them. From 'look at me' to 'I see you'. What does that look like? When you create content, think of them and what this scene or that narrative is evoking. What will they learn from you? What they want is not to passively watch how sexy you are, but for you to train them, to give them instructions, to teach them, to guide them, to be in charge, to command them. This is not being an object but the main subject. The Authority figure. How is your content already doing that. The sexy photos can still be there, they are important to already capture des attention. But what you do with that attention once you have it, is where the power dynamic is established. Positioning yourself as more than a stunning Goddess, but actually a woman who has a voice, opinions, perspective, a philosophy, a way to doing things, teaching them what you like, how you like it, why you like it, already makes them want to be that for you. You hold the attention, you hold the power, so you direct it. And for that, you want them to know you get them and you know what lives within them... that creates the desire for you to be the one exposing it. You instantly build trust. Not because you demanded it, but because you earned it: you showed them you know what you are doing. You have experience, you understand them. They are not told to come see you, they are seduced into it. They desire it. And they will work for it. This will attract better clients (real subs) and instead of you trying to get their attention, they will work to earn yours. If you want to learn more about power dynamics, building a brand as a Pro or the psychology behind BDSM, you can now access all my trainings and classes in one place for a fraction of the cost of The Dominatrix Academy. And you can reinvest the total amount towards the Program. Message me [SECRET] for the details. This offer is not available on my website.

Ms. Malissia

16,579 просмотров • 3 месяцев назад

This Should Make You Furious “A trusted teacher on a high school campus, recruiting prostitutes from her students to work for her son as their prostitutes” She “Was allowed to remain on this campus after it was brought to the school administrators attention” “This is one of the most horrendous stories we have ever had to deal with. And my history as an activist” “This was brought to their attention but the same administrators, instead of thoroughly investigating the complaint, they retaliated against the teacher who brought it to their attention when I spoke with the district the other day.” “They swore for God that they had no information about this until law enforcement called them and say, we're going to arrest Ms. Grisby today. Well, I say to Klein, Kane administrators, you are a damn lie. We got the proof.” — “Law enforcement went to her house about her daughter who was one of the victims and she came the next day and informed the school of what was happening in detail in a statement about the prostitutes ring, who's the teacher that was involved, and the young man's name that was involved. And the school district took a statement in the office with the principal and the district representative. And now you lying snakes want to tell the public that you have no idea about Miss Grisby. Oh no, that's a lie. We got it right here. We got text messages where she's texting back and forth with school administrators talking about what she brought to their attention. But yet they did nothing. The principal should be terminated.” There’s a lot more crazy info in this video, it’s absolutely disgusting this keeps happening in America. The predators are always protected. Why Does This Keep Happening In America?

Wall Street Apes

1,273,693 просмотров • 2 лет назад

Self Attention by hand ✍️ ~ 9 steps walkthrough below Self-attention is what enables LLMs to understand context. How does it work? So I drew and calculated one entirely by hand. Goal: turn four 6D features into four 3D attention weighted features, filling in every cell yourself. = 1. Given = Four feature vectors, six dimensions each, one per position. = 2. Query, key, value = Let us multiply the features by WQ, WK and WV. Queries, keys and values all come out of the same four features, and that is what the word "self" is doing in self-attention. = 3. Prepare for MatMul = We copy the queries across the top and the transposed keys down the side. Lining the two up is half the work. = 4. MatMul = Let us multiply K transpose by Q. Every cell is the dot product of one key with one query, which we use as a matching score. That works because the dot product is the numerator of cosine similarity: it is how alike two vectors are, before anyone divides by their lengths. = 5. Scale = We divide by the square root of dk, the dimension of a key vector, here 3. Without it the scores grow with the dimension and a 64-wide head would swamp the softmax. To keep the page doable in pen, the drawing approximates dividing by root 3 with halving. = 6. e to the power = Let us raise e to the power of each score. This is the first half of softmax, and the drawing uses 3 in place of e, which is close enough to do in your head. = 7. Sum = We add up each column: 16, 6, 7 and 12. = 8. Normalize = Let us divide every cell by its column sum. That gives the attention weight matrix in yellow, and each of its four columns is now a probability distribution over the four positions. The decimals are nudged as they are rounded, so every column still sums to exactly 1. = 9. MatMul = We multiply the value vectors by those weights. Each output is a blend of all four values, mixed in the proportion the attention matrix just decided, and it goes to the position-wise feed forward network in the next layer: the FFN box at the bottom of the page. The outputs: Attention weights (A), by column = [.2, .6, 0, .2], [.2, .4, .2, .2], [.4, .2, 0, .4], [.1, .7, .1, .1] Attention weighted features (Z) = [8, 2, 6], [8, 4, 4], [16, 4, 2], [4, 2, 7] The takeaway: attention is a weighted average, and everything before step 9 exists to decide the weights. Compare every position with every other, turn the scores into one distribution per position, then blend. 💾 Save this post!

Tom Yeh

27,211 просмотров • 1 месяц назад