正在加载视频...

视频加载失败

Self-Attention vs Cross-Attention by hand ✍️ interactive diagram. Open Here's some T/F questions to test your knowledge: [ ] Making X taller widens Wq, Wk and Wv but leaves Q, K and V unchanged [ ] E has to be the same length as X [ ] Q and...

14,021 次观看 • 15 天前 •via X (Twitter)

0 条评论

暂无评论

原始帖子的评论将显示在这里

相关视频

[Self-Attention] by Hand ✍️ Self-attention is what enables LLMs to understand context. How does it work? This exercise demonstrates how to calculate a 6-3 attention head by hand. Note that if we have two instances of this, we get 6-6 attention (i.e., multi-head attention, n=2). -- 𝗚𝗼𝗮𝗹 -- Transform [6D Features 🟧] to [3D Attention Weighted Features 🟦] -- 𝗪𝗮𝗹𝗸𝘁𝗵𝗿𝗼𝘂𝗴𝗵 -- [1] Given ↳ A set of 4 feature vectors (6-D): x1,x2,x3,x4 [2] Query, Key, Value ↳ Multiply features x's with linear transformation matrices WQ, WK, and WV, to obtain query vectors (q1,q2,q3,q4), key vectors (k1,k2,k3,k4), and value vectors (v1,v2,v3,v4). ↳ "Self" refers to the fact that both queries and keys are derived from the same set of features. [3] 🟪 Prepare for MatMul ↳ Copy query vectors ↳ Copy the transpose of key vectors [4] 🟪 MatMul ↳ Multiply K^T and Q ↳ This is equivalent to taking dot product between every pair of query and key vectors. ↳ The purpose is to use dot product as an estimate of the "matching score" between every key-value pair. ↳ This estimate makes sense because dot product is the numerator of Cosine Similarity between two vectors. [5] 🟨 Scale ↳ Scale each element by the square root of dk, which is the dimension of key vectors (dk=3). ↳ The purpose is to normalize the impact of the dk on matching scores, even if we scale dk to 32, 64, or 128. ↳ To simplify hand calculation, we approximate [ □/sqrt(3) ] with [ floor(□/2) ]. [6] 🟩 Softmax: e^x ↳ Raise e to the power of the number in each cell ↳ To simplify hand calculation, we approximate e^□ with 3^□. [7] 🟩 Softmax: ∑ ↳ Sum across each column [8] 🟩 Softmax: 1 / sum ↳ For each column, divide each element by the column sum ↳ The purpose is normalize each column so that the numbers sum to 1. In other words, each column is a probability distribution of attention, and we have four of them. ↳ The result is the Attention Weight Matrix (A) (yellow) [9] 🟦 MatMul ↳ Multiply the value vectors (Vs) with the Attention Weight Matrix (A) ↳ The results are the attention weighted features Zs. ↳ They are fed to the position-wise feed forward network in the next layer.

Tom Yeh

101,010 次观看 • 2 年前

Self Attention by hand ✍️ ~ 9 steps walkthrough below Self-attention is what enables LLMs to understand context. How does it work? So I drew and calculated one entirely by hand. Goal: turn four 6D features into four 3D attention weighted features, filling in every cell yourself. = 1. Given = Four feature vectors, six dimensions each, one per position. = 2. Query, key, value = Let us multiply the features by WQ, WK and WV. Queries, keys and values all come out of the same four features, and that is what the word "self" is doing in self-attention. = 3. Prepare for MatMul = We copy the queries across the top and the transposed keys down the side. Lining the two up is half the work. = 4. MatMul = Let us multiply K transpose by Q. Every cell is the dot product of one key with one query, which we use as a matching score. That works because the dot product is the numerator of cosine similarity: it is how alike two vectors are, before anyone divides by their lengths. = 5. Scale = We divide by the square root of dk, the dimension of a key vector, here 3. Without it the scores grow with the dimension and a 64-wide head would swamp the softmax. To keep the page doable in pen, the drawing approximates dividing by root 3 with halving. = 6. e to the power = Let us raise e to the power of each score. This is the first half of softmax, and the drawing uses 3 in place of e, which is close enough to do in your head. = 7. Sum = We add up each column: 16, 6, 7 and 12. = 8. Normalize = Let us divide every cell by its column sum. That gives the attention weight matrix in yellow, and each of its four columns is now a probability distribution over the four positions. The decimals are nudged as they are rounded, so every column still sums to exactly 1. = 9. MatMul = We multiply the value vectors by those weights. Each output is a blend of all four values, mixed in the proportion the attention matrix just decided, and it goes to the position-wise feed forward network in the next layer: the FFN box at the bottom of the page. The outputs: Attention weights (A), by column = [.2, .6, 0, .2], [.2, .4, .2, .2], [.4, .2, 0, .4], [.1, .7, .1, .1] Attention weighted features (Z) = [8, 2, 6], [8, 4, 4], [16, 4, 2], [4, 2, 7] The takeaway: attention is a weighted average, and everything before step 9 exists to decide the weights. Compare every position with every other, turn the scores into one distribution per position, then blend. 💾 Save this post!

Tom Yeh

27,264 次观看 • 1 个月前

Why Does Quantum Mechanics Use a Complex Wavefunction? Schrödinger’s equation doesn’t start from mystery. It starts from a very specific bet. The state of a particle is a complex field ψ(x,t), and whatever time-evolution rule we choose has to move ψ forward while preserving total probability. So the basic question is simple. What equation should ψ satisfy so that |ψ|² behaves like a conserved density, the way mass density does in fluid flow? What is ψ? Think of ψ(x,t) as an amplitude attached to the statement the particle is at position x at time t. It’s not a probability. It’s the thing you add first, and only at the end do you square it: p(x,t) = |ψ(x,t)|² Because ψ is complex, it has magnitude and phase. Write it as ψ(x,t) = r(x,t) exp(i θ(x,t)) Then r² = |ψ|² is the density, and the phase θ ends up controlling the flow through the probability current. Where does Schrödinger’s equation come from? Start with two empirical inputs that tie waves to particles: E = ħ ω p = ħ k Here ħ is Planck’s constant divided by 2π. It’s the conversion factor between frequency and energy, and between wavenumber and momentum. A plane wave with angular frequency ω and wavevector k is ψ(x,t) = A exp(i(k·x − ωt)) Now watch what derivatives do to this wave: ∂ψ/∂t = −i ω ψ ∇ψ = i k ψ ∇²ψ = −|k|² ψ Multiply by ħ and you get: i ħ ∂ψ/∂t = ħ ω ψ = E ψ −i ħ ∇ψ = ħ k ψ = p ψ −ħ² ∇²ψ = ħ² |k|² ψ = p² ψ So for plane waves, the operators Ê = i ħ ∂/∂t p̂ = −i ħ ∇ act like energy and momentum. Now bring in the classical, nonrelativistic energy bookkeeping: E = p²/(2m) + V(x) Kinetic plus potential. That’s it. Turn it into an equation for ψ by replacing E and p with the operators above: Ê ψ = (p̂²/(2m) + V) ψ Since p̂² = (−i ħ ∇)·(−i ħ ∇) = −ħ² ∇², this becomes i ħ ∂ψ/∂t = ( −ħ²/(2m) ∇² + V(x) ) ψ That’s the time-dependent Schrödinger equation. This derivation is a controlled heuristic. Match the plane-wave identities to the measured relations E = ħω and p = ħk, then impose the same energy bookkeeping you trust in classical mechanics. Why this is the right kind of rule If ψ is the state, we need a rule that preserves total probability: ∫ |ψ(x,t)|² dx = 1 Schrödinger evolution does, and you can see it by deriving a continuity equation. Let ρ(x,t) = |ψ|² = ψ*ψ. Differentiate: ∂ρ/∂t = ψ* ∂ψ/∂t + ψ ∂ψ*/∂t Use Schrödinger and its complex conjugate. The potential terms cancel, and what’s left can be rearranged into ∂ρ/∂t + ∇·j = 0 with probability current j = (ħ/(2mi)) ( ψ* ∇ψ − ψ ∇ψ* ) That’s the cleanest way to say what ψ is. |ψ|² behaves like a conserved density, the phase drives a current, and the time evolution is fixed, up to V, by combining wave relations with energy bookkeeping: i ħ ∂ψ/∂t = ( −ħ²/(2m) ∇² + V ) ψ #QuantumMechanics #SchrodingerEquation #WaveFunction #BornRule #Physics #MathematicalPhysics

Mathelirium

20,781 次观看 • 6 个月前

Lecture 2 on our Quantum Mechanics Series Schrödinger’s equation doesn’t start from mystery. It starts from a very specific bet…the state of a particle is a complex field ψ(x,t), and whatever dynamics we write down must move ψ forward in time in a way that preserves total probability. We ask a basic question…what equation should ψ satisfy so that |ψ|² behaves like a conserved density, the way mass density does in fluid flow? What is ψ? Think of ψ(x,t) as the amplitude assigned to “the particle is at position x at time t”. It’s not a probability. It’s the object you add first, and only at the end do you square p(x,t) = |ψ(x,t)|² Because ψ is complex, it has magnitude and phase. Write it in polar form ψ(x,t) = r(x,t) exp(i θ(x,t)) Then r² = |ψ|² is the density, and θ will end up controlling flow (the probability current). Where does Schrödinger’s equation come from? Start with two empirical inputs about waves and particles: E = ħ ω p = ħ k Here ħ (“h-bar”) is Planck’s constant divided by 2π. It’s the unit conversion factor between the wave description (frequency ω, wavevector k) and the particle description (energy E, momentum p). In units, ħ has units of joule-seconds, so multiplying ω (1/seconds) gives energy (joules), and multiplying k (1/meters) gives momentum (kg·m/s). It’s the number that tells you how much energy or momentum you get per unit frequency or wavenumber. A plane wave with angular frequency ω and wavevector k is ψ(x,t) = A exp(i(k·x − ω t)) Now notice what derivatives do to this wave: ∂ψ/∂t = −i ω ψ ∇ψ = i k ψ ∇²ψ = −|k|² ψ Multiply those identities by ħ: i ħ ∂ψ/∂t = ħ ω ψ = E ψ −i ħ ∇ψ = ħ k ψ = p ψ −ħ² ∇²ψ = ħ² |k|² ψ = p² ψ So for plane waves, the operators Ê = i ħ ∂/∂t p̂ = −i ħ ∇ act like energy and momentum! Now use the classical nonrelativistic energy relation: E = p²/(2m) + V(x) This is bookkeeping for a particle moving slow enough that relativity can be ignored. The term p²/(2m) is kinetic energy. If p = mv, then p²/(2m) = (m²v²)/(2m) = (1/2)mv². The term V(x) is potential energy. It depends on position because forces come from spatially varying energy. A slope in V pushes the particle. Examples: for a charged particle in an electric potential φ(x), V(x) = q φ(x). Near Earth, V(z) = mgz. The point is total energy equals kinetic plus potential. Turn that into an equation for ψ by replacing E and p with the operators above: Ê ψ = (p̂²/(2m) + V) ψ Compute p̂² = (−i ħ ∇)·(−i ħ ∇) = −ħ² ∇², so we get i ħ ∂ψ/∂t = ( −ħ²/(2m) ∇² + V(x) ) ψ That is the time-dependent Schrödinger equation. The derivation here is a controlled heuristic: we matched the plane-wave identities to the measured relations E = ħω and p = ħk, then imposed the same energy bookkeeping as classical mechanics. Why this equation is the right kind of rule If ψ is the state, we need a rule that preserves total probability: ∫ |ψ(x,t)|² dx = 1 Schrödinger evolution does. You can see it by deriving a continuity equation. Let ρ(x,t) = |ψ|² = ψ* ψ. Take a time derivative: ∂ρ/∂t = ψ* ∂ψ/∂t + ψ ∂ψ*/∂t Use Schrödinger and its complex conjugate: ∂ψ/∂t = (1/(i ħ)) ( −ħ²/(2m) ∇²ψ + Vψ ) ∂ψ*/∂t = (−1/(i ħ)) ( −ħ²/(2m) ∇²ψ* + Vψ* ) Plug in. The V terms cancel exactly, and what remains can be rearranged into a divergence: ∂ρ/∂t + ∇·j = 0 where the probability current is j = (ħ/(2mi)) ( ψ* ∇ψ − ψ ∇ψ* ) This is the best way to explain ehat ψ is: |ψ|² behaves like a conserved density, and the phase of ψ is what drives the current j. So in this series, ψ isn’t a slogan. It’s the object whose modulus squared is the density, whose phase generates flow, and whose time evolution is fixed (up to V) by matching wave relations to energy bookkeeping: i ħ ∂ψ/∂t = ( -ħ²/(2m) ∇² + V ) ψ #QuantumMechanics #SchrodingerEquation #WaveFunction #BornRule #Physics #MathematicalPhysics

Mathelirium

40,835 次观看 • 8 个月前

SORA by Hand ✍️ OpenAI’s #SORA took over the Internet when it was announced earlier this year. The technology behind Sora is the Diffusion Transformer (DiT) developed by William Peebles and Shining Xie. How does DiT work? 𝗚𝗼𝗮𝗹: Generate a video conditioned by a text prompt and a series of diffusion steps [1] Given ↳ Video ↳ Prompt: "sora is sky" ↳ Diffusion step: t = 3 [2] Video → Patches ↳ Divide all pixels in all frames into 4 spacetime patches [3] Visual Encoder: Pixels 🟨 → Latent 🟩 ↳ Multiply the patches with weights and biases, followed by ReLU ↳ The result is a latent feature vector per patch ↳ The purpose is dimension reduction from 4 (2x2x1) to 2 (2x1). ↳ In the paper, the reduction is 196,608 (256x256x3)→ 4096 (32x32x4) [4] ⬛ Add Noise ↳ Sample a noise according to the diffusion time step t. Typically, the larger the t, the smaller the noise. ↳ Add the Sampled Noise to latent features to obtain Noised Latent. ↳ The goal is to purposely add noise to a video and ask the model to guess what that noise is. ↳ This is analogous to training a language model by purposely deleting a word in a sentence and ask the model to guess what the deleted word was. [5-7] 🟪 Conditioning by Adaptive Layer Norm [5] Encode Conditions ↳ Encode "sora is sky" into a text embedding vector [0,1,-1]. ↳ Encode t = 3 to as a binary vector [1,1]. ↳ Concatenate the two vectors in to a 5D column vector. [6] Estimate Scale/Shift ↳ Multiply the combined vector with weights and biases ↳ The goal is to estimate the scale [2,-1] and shift [-1,5]. ↳ Copy the result to (X) and (+) [7] Apply Scale/Sift ↳ Scale the noised latent by [2,-1] ↳ Shifted the scaled noised latent by [-1, 5] ↳ The result is "conditioned" noise latent. [8-10] Transformer [8] Self-Attention ↳ Feed the conditioned noised latent to Query-Key function to obtain a self-attention matrix ↳ Value is omitted for simplicity [9] Attention Pooling ↳ Multiply the conditioned noised latent with the self-attention matrix ↳ The result are attention weighted features [10] Pointwise Feed Forward Network ↳ Multiply the attention weighted features with weights and biases ↳ The result is the Predicted Noise 🏋️‍♂️ 𝗧𝗿𝗮𝗶𝗻 [11] ↳ Calculate MSE loss gradients by taking the different between the Predicted Noise and the Sampled Noise (ground truth). ↳ Use the loss gradients to kick off backpropagation to update all learnable parameters (red borders) ↳ Note the visual encoder and decoder's parameters are frozen (blue borders) 🎨 𝗚𝗲𝗻𝗲𝗿𝗮𝘁𝗲 (𝗦𝗮𝗺𝗽𝗹𝗲) [12] Denoise ↳ Subtract the predicted noise from the noised latent to obtain the noise-free latent [13] Visual Decoder: Latent 🟩 → Pixels 🟨 ↳ Multiply the patches with weights and biases, followed by ReLU [14] Patches → Video ↳ Rearrange patches into a sequence of video frames.

Tom Yeh

238,303 次观看 • 2 年前

Lecture 2 of our Physics-Informed Neural Networks mini-series. In Lecture 1 we made the idea visible...a neural network isn’t predicting a PDE solution, it is the candidate function uᵩ(x,t), and the PDE residual rᵩ(x,t) is the leash that keeps it honest. Now the natural question follows: How can a neural network be punished for breaking a PDE when nobody ever handed it the true solution, and the equation itself contains derivatives like uᵩₜₜ and uᵩₓₓ? Here’s the satisfying answer: A PINN doesn’t need the true answer to be corrected. It only needs a way to measure how wrong it is according to the PDE! The network outputs uᵩ(x,t). A software called "autodiff" is used to compute the derivatives (uᵩₓ, uᵩₜ, uᵩₓₓ, …) exactly by applying the chain rule through the network. Those derivatives get dropped into the PDE to produce rᵩ(x,t). If rᵩ is big at some point, the loss spikes there, and gradient descent pushes the parameters so that rᵩ shrinks. The math breakdown We want a function u(x,t) that satisfies a PDE on a domain Ω. In this lecture we keep a concrete nonlinear example in mind, the damped sine-Gordon equation uₜₜ(x,t) + γ uₜ(x,t) − c² uₓₓ(x,t) + sin(u(x,t)) = 0. A PINN replaces the unknown function u with a neural network uᵩ(x,t), where ᵩ means all the network parameters (weights and biases). Now we build the physics residual by plugging uᵩ into the PDE rᵩ(x,t) = uᵩₜₜ(x,t) + γ uᵩₜ(x,t) − c² uᵩₓₓ(x,t) + sin(uᵩ(x,t)). If uᵩ were a true solution, rᵩ would be 0 everywhere. So we sample points (xⱼ,tⱼ) inside the domain. These are collocation points. At each one we evaluate rᵩ, and we define a physics loss L_phys(ᵩ) = meanⱼ |rᵩ(xⱼ,tⱼ)|². This is the punishment mechanism. (Punish just means: if |rᵩ| is big, L_phys is big; training updates ᵩ to make L_phys smaller. Reward means the loss drops, so those parameter changes are kept.) The key question was where the derivatives come from. Since uᵩ is built out of differentiable operations, we can compute uᵩₜ(x,t), uᵩₜₜ(x,t), uᵩₓ(x,t), uᵩₓₓ(x,t), at any input (x,t) we choose. Imagine a simple differentiable model written as a sum of nonlinear features uᵩ(x,t) = Σₖ vₖ σ( wₖx x + wₖt t + bₖ ) + b₀. Then the derivatives are just chain rule uᵩₓ(x,t) = Σₖ vₖ σ′(·) wₖx uᵩₓₓ(x,t) = Σₖ vₖ σ″(·) (wₖx)² uᵩₜ(x,t) = Σₖ vₖ σ′(·) wₖt uᵩₜₜ(x,t) = Σₖ vₖ σ″(·) (wₖt)². So rᵩ(x,t) is an explicit computable number at every (x,t). For the damped sine-Gordon example, it’s the same story, just with one extra nonlinear term: rᵩ(x,t) = [uᵩₜₜ(x,t) + γ uᵩₜ(x,t) − c² uᵩₓₓ(x,t)] + sin(uᵩ(x,t)). A real PINN is a deeper composition of these same building blocks, but it’s still just a chain rule, and autodiff is the machinery that does that bookkeeping reliably for big graphs. Then we train by gradient descent on the total loss. Even if we use only physics for the moment, the update is conceptually just ᵩ ← ᵩ − η ∇ᵩ L_phys(ᵩ), with learning rate η. In practice we also include initial/boundary conditions or data, because PDEs aren’t uniquely determined without them L(ᵩ) = L_data(ᵩ) + λ L_phys(ᵩ) + L_bc/ic(ᵩ), where L_bc/ic(ᵩ) enforces things like uᵩ(x,0) ≈ u₀(x) and uᵩₜ(x,0) ≈ v₀(x), or boundary conditions at x = ±L. So Lecture 2’s punchline is simple: the PDE becomes a training signal. We keep differentiating uᵩ, measuring rᵩ, and updating ᵩ until the residual goes quiet across Ω. #PINNs #PhysicsInformedNeuralNetworks #ScientificMachineLearning #AutoDiff #Backpropagation #PDE #DifferentialEquations #Optimization #MachineLearning #AppliedMath #ComputationalPhysics

Mathelirium

19,977 次观看 • 7 个月前

Lecture 3 of our Quantum Mechanics series. Lecture 2 gave us the one clean privilege quantum theory offers: treat ψ(x,t) as the state and ρ(x,t) = |ψ(x,t)|² as probability, because Schrödinger evolution forces ρ to obey a continuity equation. Lecture 3 is what that continuity equation is really telling you. If ρ behaves like a fluid, then the only question that matters is: What is the velocity field? Write ψ(x,t) = r(x,t) exp(i θ(x,t)). The magnitude r sets how much probability is sitting there. The phase θ sets where it tries to go. When you unpack the current j = Im(ψ* ∇ψ), it collapses to j = (ρ/m) ∇θ, which means the flow lines you draw are literally contours of phase geometry. Then the constraint that makes the picture bite: ψ has to be single-valued, so θ can’t wind by an arbitrary amount. Around any closed loop the total phase change must be 2π n, with n an integer. That’s why vortices aren’t features you add...they’re defects the math permits, in quantized units. In the render you see both layers at once...the 3D surface shows |ψ| breathing while the phase skin slides, and the 2D panel exposes the engine...current lines steering around discrete vortex charges. The math breakdown We write the state as a complex field ψ(x,t) on the plane (x in R²). The Born rule defines the probability density ρ(x,t) = |ψ(x,t)|² Schrödinger evolution (ħ = 1 units) is i ∂ψ/∂t = [ −(1/2m) ∇² + V(x,t) ] ψ Now derive conservation of probability. Start with ρ = ψ*ψ: ∂ρ/∂t = ψ* (∂ψ/∂t) + ψ (∂ψ*/∂t) Use Schrödinger and its complex conjugate: ∂ψ/∂t = (1/i) [ −(1/2m) ∇²ψ + Vψ ] ∂ψ*/∂t = (−1/i) [ −(1/2m) ∇²ψ* + Vψ* ] Substitute. The V terms cancel, and the remaining terms rearrange into the continuity equation ∂ρ/∂t + ∇·j = 0 with probability current j = (1/2mi) ( ψ* ∇ψ − ψ ∇ψ* ) = (1/m) Im(ψ* ∇ψ) So "probability density" really behaves like a conserved fluid density with flux j. Now expose the phase mechanism. Write ψ in polar form ψ(x,t) = r(x,t) exp(i θ(x,t)) Compute the gradient ∇ψ = exp(iθ) (∇r + i r ∇θ) Then ψ* ∇ψ = r (∇r + i r ∇θ) Taking the imaginary part gives Im(ψ* ∇ψ) = r² ∇θ = ρ ∇θ So the current becomes j = (ρ/m) ∇θ That’s the steering-wheel statement: Phase gradient sets the flow direction and speed (modulated by density and m). Finally, quantized vortices. Because ψ must be single-valued, going around any closed loop must return the same complex value. That forces the phase winding to be an integer multiple of 2π: ∮ ∇θ · dl = 2π n with n in Z n is the vortex charge. Vortex cores sit where ρ ≈ 0 (phase is undefined), and the current streamlines circulate around them. #QuantumMechanics #Wavefunction #SchrodingerEquation #BornRule #ProbabilityCurrent #ContinuityEquation #Phase #Vortices #TopologicalDefects #ComplexAnalysis #MathematicalPhysics #Mathematics #Physics

Mathelirium

37,998 次观看 • 8 个月前

What if Your Neural Network Was Forced to Obey Physics? Physics-Informed Neural Networks (PINNs) are neural networks trained to satisfy a differential equation by building the PDE residual directly into the loss. They emerged from a very practical problem...classical PDE pipelines can be brilliant, but they often demand heavy discretization work (meshes, stencils, stability tuning), and the method you build is usually tied to one geometry and one solver setup. A PINN flips the workflow by representing the solution itself as a smooth function uᵩ(x,t) and enforcing the physics everywhere you choose to sample the domain. People often meet PINNs in the least helpful way...via a flashy solution plot, and almost no explanation of what was enforced to get it. In this series we keep the enforcement visible. We pick a differential equation, represent the unknown solution as a flexible function, measure how well that function satisfies the equation across the domain, and train it to reduce that mismatch everywhere we sample. A normal neural net learns from labels...you give it inputs and target outputs. A PINN learns from a differential equation...you give it inputs (x,t) and it gets punished whenever its output fails the PDE. By punish we mean that the loss increases when the mismatch is large we reward it if the loss decreases as the mismatch gets smaller. The network isn’t replacing physics, it’s becoming a flexible function that is forced to satisfy the same calculus you’d impose on any candidate solution. The math breakdown: We start with a PDE we want to solve on a domain Ω. Write it as uₜ(x,t) + N(u(x,t), uₓ(x,t), uₓₓ(x,t), …) = 0 for (x,t) in Ω A PINN replaces the unknown function u with a neural network output uᵩ(x,t) Now define the physics residual by plugging uᵩ into the PDE rᵩ(x,t) = ∂uᵩ/∂t + N(uᵩ, ∂uᵩ/∂x, ∂²uᵩ/∂x², …) If uᵩ were an exact solution, we would have rᵩ(x,t) = 0 everywhere. We may also have data points (xᵢ,tᵢ,uᵢ) from measurements or a known initial condition. The training objective is just a weighted sum of squared errors L(ᵩ) = L_data(ᵩ) + λ L_phys(ᵩ) + L_bc/ic(ᵩ) with L_data(ᵩ) = meanᵢ |uᵩ(xᵢ,tᵢ) − uᵢ|² L_phys(ᵩ) = meanⱼ |rᵩ(xⱼ,tⱼ)|² where (xⱼ,tⱼ) are the collocation points in Ω L_bc/ic(ᵩ) = penalties enforcing boundary conditions and initial conditions The key technical step is that the derivatives inside rᵩ are computed by automatic differentiation ∂uᵩ/∂t, ∂uᵩ/∂x, ∂²uᵩ/∂x², … So we can differentiate the total loss L(ᵩ) with respect to ᵩ and train with gradient descent. This is the whole idea behind PINNs. Learn a function, but make the PDE part of the loss, so the network is trained to be a solution, not just a curve-fitter. In the render, the main 3D surface is the network’s current guess uᵩ(x,t), drawn as a living sheet over the (x,t) plane. Hovering above is the neural scaffold...a visible graph of feature nodes and connections. The bright tension threads are the physics residual rᵩ(x,t): each thread tethers a collocation bead on the sheet up to the scaffold, and it thickens and brightens exactly where |rᵩ| is large (color encodes the sign). As training runs, those threads go slack across the domain not because we hid the error, but because the network has actually been pushed toward rᵩ(x,t) ≈ 0. #PINNs #PhysicsInformedNeuralNetworks #ScientificMachineLearning #PDE #DifferentialEquations #Optimization #MachineLearning #AppliedMath #ComputationalPhysics

Mathelirium

17,459 次观看 • 3 个月前

Quantum Mechanics Series Lecture 4 Lecture 1 established that ρ(x,t) = |ψ(x,t)|² behaves like a conserved probability density. Lecture 2 showed what drives that flow. We also saw that writing ψ = r exp(iθ) makes the probability current proportional to the phase gradient, making it clear that phase geometry literally steers the motion. Lecture 3 then showed that the centroid of that flow can move almost classically when the packet is tight and the external potential is smooth. However, that raises yet another question. If the centroid can look classical, why does the full wave still spread, bend, split, and interfere in ways no classical particle cloud would? This is because the wave is not driven only by the external potential. It is also driven by its own curvature. Write ψ(x,t) = r(x,t) exp(iθ(x,t)) with ρ = r². Then Schrödinger’s equation gives two coupled real equations. One is the continuity equation you already know. The other looks like a Hamilton-Jacobi equation, but with one extra term: Q = −(1/2m) ∇²r / r This is the so-called Quantum Potential. It depends entirely on how the amplitude bends across space. So, the wave is being shaped not only by V(x,t), but also by the geometry of its own envelope. In the animation, the upper surface is still |ψ| and its skin is still colored by arg(ψ). The glowing threads still trace the probability current. But now a second membrane hangs underneath. That lower membrane encodes the quantum potential Q itself. The porcelain bead marks the quantum centroid. The amber bead follows a classical centroid under the same external V. When those paths separate, the lower membrane tells you why. The difference is not magic but the extra term classical mechanics does not have. The math breakdown: Start from Schrödinger evolution in units with ħ = 1: i ∂ψ/∂t = [ −(1/2m) ∇² + V(x,t) ] ψ Write the state in polar form: ψ = r exp(iθ) Then ρ = |ψ|² = r² From the imaginary part, you recover probability conservation: ∂ρ/∂t + ∇·j = 0 with j = (1/m) Im(ψ* ∇ψ) = (ρ/m) ∇θ So the local velocity field is v = j / ρ = ∇θ / m Now take the real part of Schrödinger’s equation. That gives ∂θ/∂t + |∇θ|² / (2m) + V + Q = 0 where Q = −(1/2m) ∇²r / r This is the classical Hamilton-Jacobi equation with one extra term. That extra term is what makes quantum motion locally different from classical motion. Take a gradient of that phase equation and use v = ∇θ / m. Then the flow obeys an Euler-like equation: ∂v/∂t + (v·∇)v = −(1/m) ∇(V + Q) In other words, there are really two forces in the problem. One comes from the external potential V. The other comes from the wave’s own curvature through Q. That is why Ehrenfest is only approximate. The centroid can still satisfy d⟨x⟩/dt = ⟨p⟩/m d⟨p⟩/dt = −⟨∇V⟩ but the internal shape of the packet evolves under the combined influence of V and Q. When the packet stays broad and smooth, Q is gentle and the motion looks more classical. When the packet develops sharp curvature or interference structure, Q becomes strong and the classical picture breaks down. That is what this scene is designed to show live. #QuantumMechanics #Wavefunction #SchrodingerEquation #BornRule #ProbabilityCurrent #ContinuityEquation #Phase #EhrenfestTheorem #QuantumPotential #Madelung #HamiltonJacobi #MathematicalPhysics #Mathematics #Physics

Mathelirium

20,456 次观看 • 4 个月前

ResNet by hand ✍️ ~ 10 steps walkthrough below "Deep Residual Learning for Image Recognition" (Kaiming He, CVPR 2016) is among the most cited papers in all of deep learning. Why does it matter so much? It fixed the exploding and vanishing gradients that kept deep networks from being deep, and made thousands of layers possible. How simple was the fix? An identity matrix. Goal: push three input vectors through a residual block, then through a transformer encoder block, filling in every cell yourself. = 1. Given = A mini batch of three input vectors, 3D, and the weights of the layers ahead. = 2. Linear layer = Let us multiply by the weights, add the bias, and apply ReLU so negatives become 0. Three feature vectors out. This is F(X). = 3. Concatenate = Now the trick. Stack an identity matrix beside the second layer's weights, and stack the input vectors under the features. Draw the lines between rows and columns: those are the skip connections. The identity is the residual. = 4. Linear layer + identity = We multiply the two stacked matrices. The identity carries X straight through while the weights transform it, so a single multiplication computes F(X) + X. Apply ReLU and hand it to the next block. Now watch the same trick inside a transformer, first in attention. = 5. Attention = Let us take three input vectors in 2D, compute the attention matrix, and multiply to get attention weighted vectors. = 6. Concatenate = We stack two identities this time, two residuals, which is how you get 1 + 1, and stack the input vectors with the attention weighted ones. = 7. Add = Multiply the stacked matrices. The identity adds attention to its own input, across the columns, which is how positions get combined. And again in the feed forward layer. = 8. First layer = Let us multiply by the feed forward weights and bias, then ReLU. Three feature vectors. = 9. Concatenate = Stack and link exactly as in step 3: the residual again. = 10. Second layer + identity = We multiply, apply ReLU, and pass the result to the next encoder block. This identity adds across the rows, combining features rather than positions. Takeaway: one simple "add" is what made really deep networks possible. 💾 Save this post!

Tom Yeh

18,124 次观看 • 1 个月前

Quantum mechanics has a reputation for being mystical mainly because people skip the rules and jump to interpretations. In this lecture series, we’re doing the opposite. We start from the rules, follow the algebra, and let the picture be the calculation. Classical Probability Theory combines alternatives by adding their probabilities. Quantum Theory combines them one step earlier…add complex amplitudes first, then square at the end. That swap in order is everything. Expand |a₁ + a₂|² and you don’t just get |a₁|² + |a₂|²…you get a cross-term, 2 Re(a₁ a₂*). Its sign is set by phase, so the same two contributions can reinforce or cancel. Interference is just the algebra of squaring a sum. In the 3D render, the surface height is proportional to |a(x)| (so peaks become bright bands after squaring), while the surface skin is colored by the local phase arg(a(x)). As the phase knob φ(t) is swept on path 2, the cross-term oscillates, and you literally watch the interference ridges slide across the screen. We model a detector screen with coordinates x in R² (think x = (x,y)). A quantum state assigns a complex amplitude a(x). The rule for outcomes is p(x) = |a(x)|² Now the key situation: two coherent alternatives contribute to the same outcome x. Let their amplitudes be a₁(x) and a₂(x). Quantum says a(x) = a₁(x) + a₂(x) So the probability density becomes p(x) = |a₁(x) + a₂(x)|² Expand it (this is the whole episode): p(x) = (a₁ + a₂)(a₁* + a₂*) = |a₁|² + |a₂|² + a₁ a₂* + a₁* a₂ = |a₁|² + |a₂|² + 2 Re(a₁ a₂*) That last term is the interference term. It can be positive or negative. To see phase explicitly, write each contribution in polar form: a₁(x) = r₁(x) exp(i θ₁(x)) a₂(x) = r₂(x) exp(i θ₂(x)) Then a₁ a₂* = r₁ r₂ exp(i(θ₁ − θ₂)) So the cross-term is 2 Re(a₁ a₂*) = 2 r₁ r₂ cos(θ₁(x) − θ₂(x)) That’s the fringe engine: p(x) = r₁² + r₂² + 2 r₁ r₂ cos(Δθ(x)) Now the phase knob we animate: Add a controllable phase shift φ to path 2: a₂(x) → a₂(x) exp(i φ) Then Δθ(x) → Δθ(x) − φ, so p(x; φ) = r₁² + r₂² + 2 r₁ r₂ cos(Δθ(x) − φ) As φ changes smoothly, the bright/dark pattern slides continuously. Same setup, same geometry, same magnitudes r₁,r₂, only phase changed. #QuantumMechanics #WaveInterference #ComplexAmplitudes #DoubleSlit #Physics #Mathematics

Mathelirium

81,608 次观看 • 8 个月前

Final Lecture of our Statistical Mechanics Series. Lecture 2 showed how we move from the Full Phase-Space Density ρ(q₁, …, qₙ, p₁, …, pₙ, t) to smaller statistical objects by integrating out variables we do not want to keep. That gives reduced descriptions like the One-Particle Density f₁(q₁,p₁,t) and the Two-Particle Density f₂(q₁,p₁,q₂,p₂,t) This was the simplification. Now comes the catch. If the Full Density obeys Liouville’s Equation, the reduced densities do not evolve independently. The equation for one level depends on the next one and this is referred to as the BBGKY hierarchy. The One-Particle Density depends on the Two-Particle Density. The Two-Particle Density depends on the Three-Particle Density. And the chain keeps going. That happens because particles interact. Once one particle feels the rest, one-particle information is no longer enough. Correlations enter, and the lower level is fed from above. If the full Hamiltonian is H = Σᵢ pᵢ²/(2m) + Σᵢ U(qᵢ) + (1/2) Σᵢ Σⱼ≠ᵢ Φ(qᵢ − qⱼ) then reducing the full density does not make the interaction terms disappear. It leaves behind coupling to higher-order reduced densities. So, schematically, ∂f₁/∂t + transport of one particle = interaction term involving f₂ and more generally the heirarchy is such that ∂fₛ/∂t + s-particle transport = interaction term involving fₛ₊₁ Therefore, Lecture 3 is really about the price of reduction. We simplify the description, but the information we remove comes back as coupling to higher-order correlations. So, how do you actually compute anything if every level depends on the next one? This is the so-called Closure Problem. To make the hierarchy usable, you need an extra assumption that cuts the chain. You replace the exact higher-order object by an approximation in terms of lower-order ones. The most basic example is a factorized closure at the pair level, where the exact correlated Two-Particle Density is replaced schematically by a product of One-Particle Densities: f₂(q₁,p₁,q₂,p₂,t) ≈ f₁(q₁,p₁,t) f₁(q₂,p₂,t) That approximation is not exact. It throws away part of the correlation structure. But it gives you something the raw hierarchy does not... a closed equation for the lower-level description. That is why closure matters so much. Without it, the hierarchy is exact but open. With it, the theory becomes approximate but usable. Thus, the combined point of this final Statistical Mechanics post is simple. First, reduced descriptions are not closed because interactions generate correlations across levels. Second, if you want a workable Kinetic Theory, you must close the hierarchy by approximating those higher-order correlations. It is the bridge from formal many-body mechanics to equations people can actually solve. In the render, that is exactly the story you are seeing. The first part shows the hierarchy itself: one reduced level feeding the next, with lower descriptions inheriting structure from higher ones. The second part shows the closure step where the exact correlated pair level is replaced by a factorized ansatz, and that approximation gives back a closed one-particle description. That is, the animation moves from dependence to approximation, and from approximation to solvability. #StatisticalMechanics #BBGKY #ClosureProblem #KineticTheory #PhaseSpace #ReducedDistribution #HamiltonianMechanics #MathematicalPhysics #Mathematics #Physics

Mathelirium

10,970 次观看 • 4 个月前

Graph Convolutional Network by hand ✍️ ~ 12 steps walkthrough below Graph Convolutional Networks (GCNs), introduced by Thomas Kipf and Max Welling in 2017, are the tool for data shaped like a graph: social networks, recommendations, biological networks, drug discovery, molecular chemistry. I drew and calculated a simple GCN entirely by hand. Goal: run a two-layer GCN, then a small classifier, on a five-node graph, filling in every cell yourself. 1. Given A graph of five nodes, A to E, with edges between some of them. 2. Adjacency matrix (neighbors) Put a 1 wherever two nodes share an edge, in both directions. 3. Adjacency matrix (self) Add 1s down the diagonal, one self-loop per node. That is just adding the identity matrix. 4. Messages Multiply each node's embedding by the weights and biases, then ReLU. Negatives become 0. 5. Pooling Multiply the messages by the adjacency matrix. Each node gathers the messages of its neighbours and itself. 6. Visualize Node A pools [3,0,1] + [1,0,0] = [4,0,1]. 7. Second GCN layer Messages again: weights, biases, ReLU. 8. Pooling again Pool over each node and its neighbours, once more. 9. Visualize Node C pools [1,2,4] + [1,3,5] + [0,0,1] = [2,5,10]. 10. Fully connected layer Weights, biases, ReLU. This time there are no neighbours to pool, just the node itself. 11. Linear layer One more: weights and biases. 12. Sigmoid Squash each score to a probability (≥ 3 → 1, 0 → 0.5, ≤ -3 → 0). That is the classification for each node. You have just classified every node in the graph by hand. ✍️ The outputs: A: 0 (very unlikely) B: 1 (very likely) C: 1 (very likely) D: 1 (very likely) E: 0.5 (neutral) The takeaway: a GCN layer is two parts. The top part pools each node with its neighbours through the adjacency matrix. The bottom part is an MLP that transforms each node on its own. A transformer layer has the same two parts, with an attention matrix where the adjacency matrix was. Both matrices do one job, mixing across positions: attention over tokens, adjacency over nodes. In my class I call the GCN the transformer's little cousin: a bit more stubborn, because its attention is fixed by the graph rather than computed from Q, K, and V. Draw the two side by side and the resemblance is hard to miss. 💾 Save this post! #AIbyHand #GraphNeuralNetworks #DeepLearning

Tom Yeh

16,800 次观看 • 1 个月前