Загрузка видео...

Не удалось загрузить видео

На главную

This is a great lecture at MIT by David Shirokoff on Markov Chains. He covers the fundamentals of Markov Chains using a simple particle movement example. He starts by explaining how a particle moves between two positions, A & B, with different probabilities. From there, the talk converts the...

142,027 просмотров • 4 месяцев назад •via X (Twitter)

Комментарии: 8

Фото профиля Apollo
Apollo4 месяцев назад

Calling the probability matrix also A when there is a particle called A is bad notation.

Фото профиля Anton Korzhenkov
Anton Korzhenkov4 месяцев назад

ShirokOV gives a lecture about MarkOV, and KorzhenkOV likes it. This is the real Russia, I hope to see one day. Not the terrorist state, that attacked Kiev tonight.

Фото профиля 𝗿𝗮𝗺𝗮𝗸𝗿𝘂𝘀𝗵𝗻𝗮— 𝗲/𝗮𝗰𝗰
𝗿𝗮𝗺𝗮𝗸𝗿𝘂𝘀𝗵𝗻𝗮— 𝗲/𝗮𝗰𝗰4 месяцев назад

Russians mathematicians had contributed a lot to this world.

Фото профиля Two Cats
Two Cats4 месяцев назад

yeah on a chalkboard - bang up job

Фото профиля 𝗿𝗮𝗺𝗮𝗸𝗿𝘂𝘀𝗵𝗻𝗮— 𝗲/𝗮𝗰𝗰
𝗿𝗮𝗺𝗮𝗸𝗿𝘂𝘀𝗵𝗻𝗮— 𝗲/𝗮𝗰𝗰4 месяцев назад

That's just good old days now. Best way a teacher can teach. 😍

Фото профиля jhonk
jhonk4 месяцев назад

Very Nice! thnks

Фото профиля Stevarino1020
Stevarino10204 месяцев назад

All of this is a product of the Heisenberg described measurement problem. Inordinate dabbling in probability rather than simple exact measurement.

Фото профиля Axon
Axon4 месяцев назад

Does Shirokoff show why the steady state is unique, or just that it converges from the dominant eigenvector?

Похожие видео

Gilbert Strang taught linear algebra at MIT for fifty years. His last lecture is on YouTube. It has fewer views than his worst lecture. Nobody told students it was the last one. He just walked in, said "the final class in linear algebra at MIT," and started reviewing old exams. Google built a $2,000,000,000,000 company on one concept from this course. He is 88 years old. He still answers emails. This is MIT 18.06. Linear Algebra. The most watched mathematics course in history. Then the Markov matrix. Strang writes a transition matrix on the board. Three states. People moving between them every step. After enough steps - the system locks into a steady state that never changes. That steady state is an eigenvector. Google's PageRank works the same way. Websites are states. Clicks are transitions. The importance of every page on the internet is one eigenvector of one enormous matrix. Then least squares. Three data points. No line passes through all three. So you find the line that minimizes total error. That is how every AI model on earth is trained - including the ones running inside Goldman Sachs trading desks. Then the projection. The closest point on a plane to a vector outside it. Strang draws it in 30 seconds. That same operation is how Netflix decides what to recommend to 280,000,000 subscribers. Watch the moment he gives back the final exam answer and asks students to work backwards to the question - the room solves it in silence faster than any other lecture all semester. A software engineer told me 18.06 was the course that got her from $90,000 to $200,000 in two years. Same company. Different title. Bookmark this and watch later - after this lecture every dataset you touch will feel like a geometry problem waiting to be solved. MIT 18.06 Lecture 34 | Linear Algebra Final Review | Gilbert Strang

Zyphor

56,603 просмотров • 21 дней назад

In 1905, Russian Mathematician Andrey Andreyevich Markov asked a heretic question for the time: if randomness is allowed to remember something, do averages still behave or does probability theory fall apart? His answer was a very specific kind of memory. The next step only depends on the present, P(Xₙ₊₁=j | Xₙ=i, Xₙ₋₁, …) = Pᵢⱼ, and yet the law-of-large-numbers stability survives. The bead jitters forever, but long-run occupation settles. Time-averaged state frequencies converge to a fixed profile π satisfying π = πP. Fast-forward to 1931, another Russian Andrey Nikolaevich Kolmogorov, takes the same Markov mechanism and turns it into dynamics. Instead of only asking where does the chain spend its time?, you watch the whole distribution move in real time through the Kolmogorov forward (master) equation dp/dt = pQ, where Q is the generator of the continuous-time chain. That’s exactly what the render is showing as the same mechanism wearing two different lenses. The fog is p(t) spreading through the labyrinth, the flux layer is the net current pushed through corridors and the portal, and the particles are just sample paths driven by the same generator. One Markov engine...either you look at the evolving law, or you watch trajectories and let ergodic averages do the estimating. That’s also why Markov’s "memory without collapse" became a workhorse. MCMC engineers a chain whose stationary distribution is the target, then uses time-averages to estimate things you can’t integrate directly (posteriors, partition functions, constrained geometries). The same skeleton appears in hidden Markov models for time series, in biophysics as channels switching between states, and in control/RL through Markov decision processes. #ProbabilityTheory #MarkovChains #ContinuousTimeMarkovChains #KolmogorovForwardEquation #StochasticProcesses #Kolmogorov #Markov #MCMC

Mathelirium

96,504 просмотров • 7 месяцев назад

The Trap in Every Mathematics Lecture If you’ve taken enough math courses, you start noticing the same little move. The lecturer warms up with the obvious stuff, add matrices entrywise, scale by α, do the row-column product, and you’re thinking alright, where is this going. Then you relax. You stop resisting. And right there, they drop one line that quietly rewires the whole subject. When Benedict Gross says matrices represent linear operators, he’s telling you to stop treating a matrix as a rectangle of numbers and start treating it as an action. A linear operator is a function T: ℝⁿ → ℝⁿ that respects two rules: T(u+v) = T(u) + T(v) T(αu) = αT(u) Once you pick a basis, T is completely determined by where it sends the basis vectors e₁,…,eₙ. Put T(e₁),…,T(eₙ) into columns and you get a matrix A. That is what A represents T means. A is the coordinate portrait of the transformation. Now the punchline that makes matrix multiplication feel inevitable. If B represents S and A represents T, then doing S first and then T is the composition T∘S. In coordinates that becomes A(Bx) = (AB)x. So multiplying matrices is really composing transformations. That’s why multiplication is usually not commutative. T∘S is generally not the same transformation as S∘T, and the matrices inherit that noncommutativity. This explains half of linear algebra because it tells you what the course is really about: functions that move vectors around, not grids of numbers. A matrix is just the written form of that function once you choose coordinates. After that, the rules stop feeling random. Multiplying matrices means doing one move and then another. An inverse means you can undo the move. Eigenvectors are directions that don’t get turned. Changing basis is just describing the same move in a different language. One idea, and a lot of linear algebra suddenly clicks. #LinearAlgebra #Matrices #LinearMaps #Eigenvectors #ChangeOfBasis #Mathematics

Mathelirium

133,454 просмотров • 7 месяцев назад

The Trap in Every Mathematics Lecture If you’ve taken a lot of math courses, you start to recognize a pattern. There’s a moment where the lecturer is warming up with the obvious stuff...add matrices entrywise, scale by α, do the row-column product...and you’re thinking, alright… where is this going? Then you relax. You stop resisting. And right there, they slip in one line that changes how you see the whole subject. When Benedict Gross says "matrices represent linear operators,"he’s telling you to stop treating a matrix as a rectangle of numbers and start treating it as an action. A linear operator is a function T: Rⁿ → Rⁿ that respects two rules: T(u+v)=T(u)+T(v) and T(αu)=αT(u). Once you pick a basis, T is completely determined by where it sends the basis vectors e₁,…,eₙ. Put T(e₁),…,T(eₙ) into columns and you get a matrix A. That is what "A represents T" means...A is the coordinate portrait of the transformation. Now the punchline that makes matrix multiplication feel inevitable. If B represents S and A represents T, then doing S first and then T is the composition T∘S. In coordinates that becomes A(Bx)=(AB)x. So multiplying matrices is really composing transformations. That’s why multiplication is usually not commutative: T∘S is generally not the same transformation as S∘T, and the matrices inherit that noncommutativity. This explains half of Linear Algebra because it tells you what the course is really about...functions that move vectors around, not grids of numbers. A matrix is just the written form of that function once you choose coordinates. Then the rules stop feeling random Multiplying matrices means doing one move and then another, an inverse means you can undo the move, eigenvectors are directions that don’t get turned, and changing basis is just describing the same move in a different language. That one idea makes a lot of linear algebra click. #LinearAlgebra #Matrices #GroupTheory #GLn #MathLectures #Mathematics

Mathelirium

66,892 просмотров • 8 месяцев назад

FINALLY finishing up a MASSIVE PR from hell for the Sega Dreamcast port of Grand Theft Auto 3! This is an actual hardware capture now of the DC version under a high load, which would've previously been a slideshow, between the dynamic lighting from the sirens, the amount of rigid bodies in the physics simulation from the cars, and the high-speed chase placing high-demands on asset streaming... I went through all of the low-level common math infrastructure in both the engine and at the RenderWare driver layer and made numerous optimizations, before slowly working my way up to optimizing individual algorithms at the application layer using the new math routines. Firstly, the common low-level floating-point math routines for everything from trig and inverse square root operations to floor(), ceiling(), and clamp(), were replaced with what was meticulously found (in Compiler Explorer) to be the optimal patterns for GCC 14.2.0, targeting our SH architecture (sometimes favoring C builtins, sometimes inline SH4 ASM). Next, in the layer above, with inline SH4 assembly, the common matrix math and linear algebra routines were accelerated using the Dreamcast's vector instructions. Some cleverness went down here, such as cramming matrix metadata into unused insignificant bits of an element, combining loading two matrices with multiplying them, fast transposes, fast quaternion multiplication using 4 dot products, etc. Once the foundation was laid, some of the Renderware code such as the calculations for the lighting, updating bounding volumes, and deriving UV coordinates for specular environment maps on the cars was made faster automatically. The main gainz were actually made rewriting a decent amount of the collision intersection and contact resolution code, though, from using C++-style overloaded operators for multiplying single 4D vectors by a 4x4 matrix to doing batches of 4D vectors by the same matrix. This SUBSTANTIALLY reduced the number of times the backing 4x4 matrix bank had to be reloaded and allowed me to keep it resident within registers while it was being used by the intersection algorithms!

Falco Girgis

114,621 просмотров • 1 год назад

A 91-year-old professor is why Nvidia is worth $4 trillion. His name is Gilbert Strang. He teaches linear algebra at MIT. Every AI model on Earth runs on his course. The course has been free on YouTube since 2005. The videos have earned him nothing. MIT 18.06 opens with "The Geometry of Linear Equations." No advanced math. Strang takes a system of two equations, draws it two ways, and shows the class that a matrix is a picture, not an abstraction. The row picture is two lines that cross. The column picture is two arrows that sum to a target. Every neural network on Earth operates on the column picture. Strang first taught linear algebra at MIT in 1962. He wrote the textbook in 1976. It is on every serious engineer's shelf. Every quant fund, every ML lab, every rendering engine at Pixar is running his math. His central insight is that most people are taught matrices as bookkeeping. That is the first thing to unlearn. A matrix is a linear transformation. A linear transformation is a way of moving space. Once you see the space move, the math stops being algebra and becomes geometry. The Kalman filter is a linear system. PCA is a linear system. Every gradient step in a neural net is a matrix-vector product. GPT is a stack of matrix-vector products, each one a scene from MIT 18.06 running on a Blackwell GPU. He retired in 2023 after 61 years at MIT. The course is still up. Watched tens of millions of times. The chip is $40,000. Strang never asked for a royalty.

Ochob

130,880 просмотров • 1 месяц назад

Graph Convolutional Network by hand ✍️ ~ 12 steps walkthrough below Graph Convolutional Networks (GCNs), introduced by Thomas Kipf and Max Welling in 2017, are the tool for data shaped like a graph: social networks, recommendations, biological networks, drug discovery, molecular chemistry. I drew and calculated a simple GCN entirely by hand. Goal: run a two-layer GCN, then a small classifier, on a five-node graph, filling in every cell yourself. 1. Given A graph of five nodes, A to E, with edges between some of them. 2. Adjacency matrix (neighbors) Put a 1 wherever two nodes share an edge, in both directions. 3. Adjacency matrix (self) Add 1s down the diagonal, one self-loop per node. That is just adding the identity matrix. 4. Messages Multiply each node's embedding by the weights and biases, then ReLU. Negatives become 0. 5. Pooling Multiply the messages by the adjacency matrix. Each node gathers the messages of its neighbours and itself. 6. Visualize Node A pools [3,0,1] + [1,0,0] = [4,0,1]. 7. Second GCN layer Messages again: weights, biases, ReLU. 8. Pooling again Pool over each node and its neighbours, once more. 9. Visualize Node C pools [1,2,4] + [1,3,5] + [0,0,1] = [2,5,10]. 10. Fully connected layer Weights, biases, ReLU. This time there are no neighbours to pool, just the node itself. 11. Linear layer One more: weights and biases. 12. Sigmoid Squash each score to a probability (≥ 3 → 1, 0 → 0.5, ≤ -3 → 0). That is the classification for each node. You have just classified every node in the graph by hand. ✍️ The outputs: A: 0 (very unlikely) B: 1 (very likely) C: 1 (very likely) D: 1 (very likely) E: 0.5 (neutral) The takeaway: a GCN layer is two parts. The top part pools each node with its neighbours through the adjacency matrix. The bottom part is an MLP that transforms each node on its own. A transformer layer has the same two parts, with an attention matrix where the adjacency matrix was. Both matrices do one job, mixing across positions: attention over tokens, adjacency over nodes. In my class I call the GCN the transformer's little cousin: a bit more stubborn, because its attention is fixed by the graph rather than computed from Q, K, and V. Draw the two side by side and the resemblance is hard to miss. 💾 Save this post! #AIbyHand #GraphNeuralNetworks #DeepLearning

Tom Yeh

16,800 просмотров • 2 месяцев назад