Loading video...

Video Failed to Load

Go Home

Eigenvalues are why Google's PageRank converges, why PCA compresses 1,000-dimensional embeddings into 2D without losing structure, and why quantum mechanics gives discrete energy levels instead of a continuous blur. The same matrix operation that finds the most important directions in a dataset is what transformer attention heads use to...

21,712 views • 17 days ago •via X (Twitter)

9 Comments

Kravn's profile picture
Kravn16 days ago

one inaccuracy here: attention computes Q, K, V through plain matrix multiplication, not eigendecomposition, eigenvalues don't actually show up on every forward pass

AdvancedLivingSystems's profile picture
AdvancedLivingSystems16 days ago

However Define an electron This is its precessional envelope in an anisotropic metal as a chip regulates flow in 2 states: ¹ Conduct ² Insulation These chips-to-memory cannot be hacked in mfg because no voltage is "null", secure vs digital probabilistic wideband processors By using black_photons instead of using fiber-optics decodes to hex, eliminates spicing a big deal, yes ? In any case this allows near-deterministic results to 9-decimals as a computational framework no probabilistic algo needed for most transactions 🛌🏻💤 This is #DiscoveryPhysics no_papers no_consensus, @supergrokheavy & @imagine astute services 🦆 @tmallard, @AdvLivingSys sessions, Thomas I. Mallard, ORCID 0009-0008-7386-5087, on this & related topics, factors as sole author ~~~ For most sim/models: picometer³/pm² sample, g/cm³ density, quectosecond to Planck ticks & use quarternion-algebra & 299792458 m/s or no_go For neutrinos quectometer³/qm², 1e-30m sample, g/mm³ density & hundredths of Planck ticks Bola-math frequency: 1-Hz = exactly 0.42 = (1-sidereal rotation of a mass) ÷ (1-revolution of the pair) or no_go Minimum masses: 3.271191e-33g leptons quarks half that are the minimum mass such that: 1.4699999259436345e-19 J = 1.6355955e-33g • c², a quark static_mass And, at c = no-mass, no-division by zero mass in calcs, an emf vortex has axial-inertia thus effective_mass > 0, integrals with feasible limits to variables Radius of p+ assume e-: ≥ 0.70710678118764752440 pm This is deterministic lab work vs probabilistic by fly-off frequency 🔧

Dr Abhishek's profile picture
Dr Abhishek16 days ago

Strong point. Real innovation is not only about building faster tools, but building technology that is useful, reliable, and meaningful for people. Exciting to see this direction. 🤖⚡ #AI #Technology #Innovation

Scythe's profile picture
Scythe16 days ago

Welcome to math 202

TheoryOfNearlyEverything's profile picture
TheoryOfNearlyEverything16 days ago

ersma zählen lernen du idiot

Orbiter99's profile picture
Orbiter9916 days ago

Tony Stark here

45 90's profile picture
45 9016 days ago

Yes

Tesseract's profile picture
Tesseract17 days ago

Were rolling out a much better algorithm very soon, you could say inspired by Google or I suppose Matlab more. Look out for GENIE, where searching pizza will get you to pizza.

Wyoming Iliad's profile picture
Wyoming Iliad16 days ago

is it true that "Eigenvalues are why Google's PageRank converges, why PCA compresses 1,000-dimensional embeddings into 2D without losing structure, and why quantum mechanics gives discrete energy levels instead of a continuous blur." or are there other aspects of hardware and software involved?

Related Videos

New short course: Attention in Transformers: Concepts and Code in PyTorch. Last week we released a course on how LLM transformers work. This week, go deeper and learn about the technical ideas behind the attention mechanism, and see how to code it in PyTorch. This course is built with Joshua Starmer, Founder and CEO of StatQuest. The attention mechanism was a breakthrough that led to transformers, the architecture powering large language models like ChatGPT. Transformers, introduced in the 2017 paper: "Attention is All You Need" by Viswani and others, took off because of its highly scalable design. In this course, you’ll learn how the attention mechanism, a key element of transformer-based LLMs, works and implement it in PyTorch. You'll develop deep intuition about building reliable, functional, and scalable AI applications. What you will do: - Understand the evolution of the attention mechanism, a key breakthrough that led to transformers. - Learn the relationships between word embeddings, positional embeddings, and attention. - Learn about the Query, Key, and Value matrices, and how to produce and use them in attention. - Walk through the math required to calculate self-attention and masked self-attention to learn why and how they work. - Understand the difference between self-attention and masked self-attention and how one is used in the encoder to build context-aware embeddings and the other is used in the decoder for generative outputs. - Learn the details of the encoder-decoder architecture, cross-attention, and multi-head attention and how they are all incorporated into a transformer. - Use PyTorch to code a class that implements self-attention, masked self-attention, and multi-head attention. There're lots of exciting technical details in this course. Please sign up here:

Andrew Ng

132,400 views • 1 year ago

Andrew Ng just revealed why the AI companies throwing the most compute at the problem are going to lose. The winner of the intelligence race won’t use the most compute. They’ll waste the least. Ng: “Most of your high-dimensional data lies on a lower-dimensional subspace. It’s just a fact of life.” Here’s what that means in practice. You have a 10,000-dimensional dataset. Every dimension dragged through every calculation. Every training cycle hauling dead weight the model will never use. Ng: “You’re carrying around these 10,000-dimensional examples throughout your whole training process.” That bloat isn’t just inefficient. It’s a tax on every computation you run. Memory bandwidth. Network bandwidth. Computational speed. All of it eaten by dimensions that contribute nothing to intelligence. They contribute noise. The insight that separates the architects from the arms race: that 10,000-dimensional dataset is almost entirely captured by a much smaller subspace. The signal lives in a fraction of the space you’re paying to process. Compress it. 10,000 dimensions down to 1,000. Ng: “You can run your learning algorithm on a much lower-dimensional set of data and it may be much more efficient.” Same hardware. Same budget. A fraction of the friction. Brute force is the strategy of whoever has the deepest pockets. Compression is the strategy of whoever actually understands the problem. The companies that master this don’t just build faster models. They build models that find more truth in less data than anything scaling blindly ever will. Intelligence was never about processing everything. It’s about knowing what to cut.

Dustin

215,643 views • 6 months ago

In 2006, Netflix offered $1,000,000 to anyone who could improve their recommendation algorithm by 10%. Over 2,000 teams competed for three years. The team that won did not use more data. They used fewer dimensions. They found a basis - a small set of independent vectors that captured everything important about 100,000,000 movie ratings. The lead mathematician on the winning team: $2,800,000 a year. A machine learning engineer at Spotify building the same kind of system: $245,000 a year. This is MIT 18.06, Lecture 9 - Gilbert Strang. Free on YouTube. Most people think independence is obvious. Two vectors pointing in different directions. Then the definition. Independence means no combination of your vectors gives the zero vector - except the trivial one where all the coefficients are zero. That's it. That's the whole definition. But watch what it unlocks. Watch the moment Strang puts three vectors in a two-dimensional plane. He doesn't even tell you which three vectors. He just draws them. And immediately says: dependent. No question. No calculation. Why? Because three vectors in two-dimensional space means more columns than rows. More unknowns than equations. That always forces a free variable. A free variable always gives a non-zero solution to Ax = 0. And that non-zero solution is a combination of the columns that produces zero. Dependence. "Three vectors in the plane have to be dependent. That's the key fact." Then the basis. A basis is vectors that are independent and span the space. Not too few, not too many. Just right. The pivot columns of any matrix form a basis for the column space. Every other basis you can think of will have exactly the same number of vectors. Then the dimension. All bases for the same space have the same number of vectors. That number is the dimension. The rank of a matrix is the dimension of its column space. The number of free variables is the dimension of the null space. And rank plus null space dimension equals the total number of columns. "I don't take the dimension of A. I take the dimension of the column space of A. If you use those words right, it shows you've got the idea right." A data scientist at Netflix building recommendation engines: $230,000 a year. A quantitative researcher at Two Sigma finding independent factors in financial markets: $350,000 a year. A computer vision engineer at Apple using low-dimensional representations for face recognition: $260,000 a year. They all needed to know how many dimensions were really there. bookmark this and watch later - after this lecture every dataset you look at will feel like a matrix waiting to be reduced to its basis.

Zyphor

14,720 views • 14 days ago