Загрузка видео...

Не удалось загрузить видео

На главную

This is why we need to gate-keep science. Two pseudo-intellectuals thinking they discovered something deep, conflating >social attention >transformers attention >quantum physics observer (attention) These have nothing in common, other than the ambiguity of English language. Naked ladies on Instagram have nothing to do with a weighted average followed...

119,551 просмотров • 2 месяцев назад •via X (Twitter)

Комментарии: 0

Нет доступных комментариев

Здесь появятся комментарии из оригинального поста

Похожие видео

New short course: Attention in Transformers: Concepts and Code in PyTorch. Last week we released a course on how LLM transformers work. This week, go deeper and learn about the technical ideas behind the attention mechanism, and see how to code it in PyTorch. This course is built with Joshua Starmer, Founder and CEO of StatQuest. The attention mechanism was a breakthrough that led to transformers, the architecture powering large language models like ChatGPT. Transformers, introduced in the 2017 paper: "Attention is All You Need" by Viswani and others, took off because of its highly scalable design. In this course, you’ll learn how the attention mechanism, a key element of transformer-based LLMs, works and implement it in PyTorch. You'll develop deep intuition about building reliable, functional, and scalable AI applications. What you will do: - Understand the evolution of the attention mechanism, a key breakthrough that led to transformers. - Learn the relationships between word embeddings, positional embeddings, and attention. - Learn about the Query, Key, and Value matrices, and how to produce and use them in attention. - Walk through the math required to calculate self-attention and masked self-attention to learn why and how they work. - Understand the difference between self-attention and masked self-attention and how one is used in the encoder to build context-aware embeddings and the other is used in the decoder for generative outputs. - Learn the details of the encoder-decoder architecture, cross-attention, and multi-head attention and how they are all incorporated into a transformer. - Use PyTorch to code a class that implements self-attention, masked self-attention, and multi-head attention. There're lots of exciting technical details in this course. Please sign up here:

Andrew Ng

132,285 просмотров • 1 год назад

[Self-Attention] by Hand ✍️ Self-attention is what enables LLMs to understand context. How does it work? This exercise demonstrates how to calculate a 6-3 attention head by hand. Note that if we have two instances of this, we get 6-6 attention (i.e., multi-head attention, n=2). -- 𝗚𝗼𝗮𝗹 -- Transform [6D Features 🟧] to [3D Attention Weighted Features 🟦] -- 𝗪𝗮𝗹𝗸𝘁𝗵𝗿𝗼𝘂𝗴𝗵 -- [1] Given ↳ A set of 4 feature vectors (6-D): x1,x2,x3,x4 [2] Query, Key, Value ↳ Multiply features x's with linear transformation matrices WQ, WK, and WV, to obtain query vectors (q1,q2,q3,q4), key vectors (k1,k2,k3,k4), and value vectors (v1,v2,v3,v4). ↳ "Self" refers to the fact that both queries and keys are derived from the same set of features. [3] 🟪 Prepare for MatMul ↳ Copy query vectors ↳ Copy the transpose of key vectors [4] 🟪 MatMul ↳ Multiply K^T and Q ↳ This is equivalent to taking dot product between every pair of query and key vectors. ↳ The purpose is to use dot product as an estimate of the "matching score" between every key-value pair. ↳ This estimate makes sense because dot product is the numerator of Cosine Similarity between two vectors. [5] 🟨 Scale ↳ Scale each element by the square root of dk, which is the dimension of key vectors (dk=3). ↳ The purpose is to normalize the impact of the dk on matching scores, even if we scale dk to 32, 64, or 128. ↳ To simplify hand calculation, we approximate [ □/sqrt(3) ] with [ floor(□/2) ]. [6] 🟩 Softmax: e^x ↳ Raise e to the power of the number in each cell ↳ To simplify hand calculation, we approximate e^□ with 3^□. [7] 🟩 Softmax: ∑ ↳ Sum across each column [8] 🟩 Softmax: 1 / sum ↳ For each column, divide each element by the column sum ↳ The purpose is normalize each column so that the numbers sum to 1. In other words, each column is a probability distribution of attention, and we have four of them. ↳ The result is the Attention Weight Matrix (A) (yellow) [9] 🟦 MatMul ↳ Multiply the value vectors (Vs) with the Attention Weight Matrix (A) ↳ The results are the attention weighted features Zs. ↳ They are fed to the position-wise feed forward network in the next layer.

Tom Yeh

101,010 просмотров • 2 лет назад

"When we decline conversation with those we might currently disagree with, we condemn ourselves to forever disagree." Leading into tomorrow's protests, some have instructed protesters not to engage in conversation with those on the 'other side'. If anyone believes that these identity politic divides somehow resolve themselves by forcing the 'other side' to accept your position... well... world history has already shown us how well that works out. We gain traction in this division only by first humanizing those we disagree with. By truly engaging with and understanding both their needs and their fears. Once we've got that down, we can address both our own concerns AND theirs more easily than we may have thought. If you're trans and you don't think anything gender critical people say has merit, then I have a question for you: Have you ever really talked to a gender critical individual? Not yelled at them, but sat down over a cup coffee to understand their fears? If you've done that... great! Keep going... and 6 or 12 months later you'll have talked to enough people to REALLY understand the root of where they are coming from. And on the other side... if you have concerns about schools, curriculum, childhood transition, pronounce, prisons, sports, washrooms... If you have these concerns, have you taken the time to really get into both the headspace and heart-space of a transgender person? Not to mock them on Twitter, but to engage in good old-fashion community? I've spent the last two years having thousands of conversations with the individuals on both ends of this identity politic divide. And I've never been more confident that a workable solution exists. But to get there.... we need to talk. All of us. Tomorrow, I hope you consider extending goodwill to those on the other side of the protest. That will make a much bigger difference than yelling into the abyss about how wrong they are. Maybe they are wrong... but by showing the worst of ourselves to others we only strengthen their resolve that our position is inimical.

Julia Malott

129,546 просмотров • 2 лет назад

In light of the names Anneke Lucas mentioned, like the former Canadian prime minister who is dead now, but is the father of the current one. It is being shared all over the internet now. Canadian accounts are posting this, and it is getting lots of attention. This is also a wake-up call for Canada. A lot of people in Canada are commenting on social media that they are shocked by this guy being named. But actually he was already mentioned by Cathy O'Brien in 1995, as being involved in these networks! See the video below. So there is corroboration between these witnesses. A lot of these former and now deceased prime ministers she mentions were, at that time that she was in it in the 70s, involved in this. Also the Belgian one that Anneke names, and was also named by other witnesses. Which tells us, like these witnesses say, this was a global network. So the truth was already out there a bit, but it is coming out more for a bigger public. We will have more revelations coming out in 2025 about these dark dealing that have been going on at the highest levels, I'm sure. This is going to be a shock for a lot of people that are new to this information, because it is reaching a bigger public. Some will choose not to believe it, do no research, and look no further. But a lot more people will wake up to this, it is unstoppable. Of course, we have to check in with our own vibration. If it gets too heavy, take some time out and go for a walk in nature. Whilst revelations are important for the truth to come out, and confronting hard information is part of that, we have to check with our selves how this information is affecting us, and sometimes we need a time off. It is always most important to maintain a healthy balance between checking in with this information and our own peace of mind. #Epstein #PBDpodcast

Jens Patteeuw

20,426 просмотров • 1 год назад