Loading video...

Video Failed to Load

Go Home

Zhilin at GTC: Introducing Attention Residuals Learning selective memory, rather than mechanically accumulating everything, is the beauty of attention. Many of you have probably read Attention Is All You Need, the 2017 Transformer paper that brought “human-like” attention into the model’s field of view. From that point on, models...

115,162 views • 5 months ago •via X (Twitter)

0 Comments

No comments available

Comments from the original post will appear here

Related Videos

New short course: Attention in Transformers: Concepts and Code in PyTorch. Last week we released a course on how LLM transformers work. This week, go deeper and learn about the technical ideas behind the attention mechanism, and see how to code it in PyTorch. This course is built with Joshua Starmer, Founder and CEO of StatQuest. The attention mechanism was a breakthrough that led to transformers, the architecture powering large language models like ChatGPT. Transformers, introduced in the 2017 paper: "Attention is All You Need" by Viswani and others, took off because of its highly scalable design. In this course, you’ll learn how the attention mechanism, a key element of transformer-based LLMs, works and implement it in PyTorch. You'll develop deep intuition about building reliable, functional, and scalable AI applications. What you will do: - Understand the evolution of the attention mechanism, a key breakthrough that led to transformers. - Learn the relationships between word embeddings, positional embeddings, and attention. - Learn about the Query, Key, and Value matrices, and how to produce and use them in attention. - Walk through the math required to calculate self-attention and masked self-attention to learn why and how they work. - Understand the difference between self-attention and masked self-attention and how one is used in the encoder to build context-aware embeddings and the other is used in the decoder for generative outputs. - Learn the details of the encoder-decoder architecture, cross-attention, and multi-head attention and how they are all incorporated into a transformer. - Use PyTorch to code a class that implements self-attention, masked self-attention, and multi-head attention. There're lots of exciting technical details in this course. Please sign up here:

Andrew Ng

132,400 views • 1 year ago

I hear so often from the Dommes I work with that they struggle with people online fetichizing them and simply seeing them for how sexy and beautiful they are. They project their fantasies and their desires onto you. That stops immediately once you move the attention from you to them. From 'look at me' to 'I see you'. What does that look like? When you create content, think of them and what this scene or that narrative is evoking. What will they learn from you? What they want is not to passively watch how sexy you are, but for you to train them, to give them instructions, to teach them, to guide them, to be in charge, to command them. This is not being an object but the main subject. The Authority figure. How is your content already doing that. The sexy photos can still be there, they are important to already capture des attention. But what you do with that attention once you have it, is where the power dynamic is established. Positioning yourself as more than a stunning Goddess, but actually a woman who has a voice, opinions, perspective, a philosophy, a way to doing things, teaching them what you like, how you like it, why you like it, already makes them want to be that for you. You hold the attention, you hold the power, so you direct it. And for that, you want them to know you get them and you know what lives within them... that creates the desire for you to be the one exposing it. You instantly build trust. Not because you demanded it, but because you earned it: you showed them you know what you are doing. You have experience, you understand them. They are not told to come see you, they are seduced into it. They desire it. And they will work for it. This will attract better clients (real subs) and instead of you trying to get their attention, they will work to earn yours. If you want to learn more about power dynamics, building a brand as a Pro or the psychology behind BDSM, you can now access all my trainings and classes in one place for a fraction of the cost of The Dominatrix Academy. And you can reinvest the total amount towards the Program. Message me [SECRET] for the details. This offer is not available on my website.

Ms. Malissia

17,098 views • 4 months ago

I am a physicist and I'm profoundly opposed to any idea of non-physical explanations that contradict physics. So that's a no-no and really doesn't make sense. However, there are ways in which both emergent properties such as minds and life and so on have an effect. And as you said, also abstractions. Now the fact that the theory of good explanations led to the idea that abstractions are real things was slightly surprising to me. I wasn't expecting the link, at least wasn't expecting it to be so strong as it is. But the thing is, if you think about how to explain events, physical events like a footprint on the moon, how do you explain how that happened? Well, it happened because of human ideas, of science. And human ideas, you could say in this reductionist sense that as you rightly say is the prevailing mode of explanation and the prevailing idea is to look down on other modes of explanation, that those ideas are nothing more than configurations of atoms. So some physicists, some rocket scientists put their brain into certain configurations of atoms and those atoms then acted on other atoms which then ended up making a footprint on the moon. Now what that misses is the explanation of why certain configurations of atoms put footprints on the moon while others, the overwhelming majority of configurations that human brains, even human brains have been put into in history, do not have that effect. And it's because there's a certain type of information. And this information can't in my view be reduced to statements about atoms because if you think about what that information does, it is in brains but the same information then gets transferred into, let's say, sound waves in air and then it gets transferred into ink on paper and then it gets transferred into magnetic domains inside a computer which then control a machine that instantiates those ideas in bits of steel and silicon and so on and so on. There's an immense chain of instantiations of the same information. And it's only special kinds of information that have this property that they are preserved and instantiated in successive physical modes. So what is being transmitted, what is having the causal effect is not the atoms but the fact that the atoms instantiate certain kinds of information and not other kinds. So therefore it is the information that is having the causal effect. If a particular instantiation of that information were damaged, then processes would come along to fix it, whether or not they could fix the physical instantiation. For example, if the computer goes wrong, then we don't use the corrupted information. We go back and rescue the information from a different computer and we throw away the atoms that at one point instantiated it. So the information causes itself to remain in existence. Now I think there's no way out of that mode of explanation. And if explanation is going to be the fundamental thing about our criterion, for example, about what is or isn't real, then we have to say that information and this particular kind which we call knowledge is real and really does cause things. David Deutsch

Deutsch Explains

45,735 views • 2 years ago