Загрузка видео...

Не удалось загрузить видео

На главную

Once you’ve accepted that Brownian motion doesn’t have a classical dX/dt but does have this rigid quadratic variation, you’re ready for the first real Itô vs chain rule punch: take X_t = B_t and look at f(x) = x². In ordinary calculus you’d write d(B_t²) = 2 B_t dB_t...

31,380 просмотров • 9 месяцев назад •via X (Twitter)

Комментарии: 0

Нет доступных комментариев

Здесь появятся комментарии из оригинального поста

Похожие видео

MIT FILMED A PROFESSOR WHOSE EXPLANATIONS ARE SO CLEAR AND INTUITIVE THAT THE MOST COMPLEX CALCULUS ON WALL STREET BECOMES OBVIOUS - AND IT PROVES WHY NEWTON'S VERSION GIVES THE WRONG ANSWER FOR EVERY OPTION PRICE This is an MIT lecture on Itô calculus, course 18.S096. The professor opens with one equation that breaks classical calculus - dB squared equals dt. In normal calculus, dt squared vanishes when you take limits. But the quadratic variation of Brownian motion forces the second derivative term to survive. This one fact changes every formula in calculus. Then Itô's lemma. In classical calculus, if you apply a smooth function to a path, you only need the first derivative term. In Itô calculus you need an extra term - half the second derivative times dt. He derives it in 10 lines of Taylor expansion, showing exactly which terms survive and which vanish. The correction term is not a small adjustment. It is the difference between the right answer and the wrong one. Then geometric Brownian motion. The natural guess for modeling stock prices is to take the exponential of a Brownian motion. It fails - the Itô correction introduces an unwanted drift term that makes the expected value grow over time. The fix is to subtract exactly half of sigma squared from the exponent. This is the formula that runs every Black-Scholes calculation on earth. Then adapted processes and Itô integrals. In classical integration it doesn't matter which point in each interval you use to form Riemann sums - the limit is always the same. In Itô calculus left endpoints and right endpoints give different answers. Itô's choice of left endpoints is not arbitrary - it encodes the fundamental constraint that financial decisions must be made using only past information. Watch the moment he explains the Girsanov theorem - that a Brownian motion with drift and a Brownian motion without drift are equivalent probability measures. Two processes whose paths look completely different in the long run can be converted into each other by multiplication. This is how quants transform non-martingale stock price processes into martingales for pricing. A derivatives trader I know rewatched this lecture before his first day at a quantitative hedge fund. Said it was the first time Itô's lemma felt like a theorem with a reason rather than a formula to memorize. Free on YouTube, MIT OpenCourseWare, Creative Commons license. bookmark this and watch later - after this lecture every option price you see will feel like a solution to a stochastic differential equation waiting to be written down

Zyphor

47,224 просмотров • 2 месяцев назад

A Japanese mathematician published a result in 1944 that nobody understood for twenty years. Today it runs inside every options desk on Wall Street. Goldman pays $400K to quants who can derive it from scratch and explain why classical calculus gives the wrong answer without it. His name is Choongbum Lee. MIT, 18.S096, Topics in Mathematics with Applications in Finance. The course that Wall Street watches. This is lecture 17. It derives Ito's Lemma from scratch. He opens with the problem nobody in classical calculus can solve. Then the foundation. Brownian motion is the limit of a random walk taken to infinity. Each trade pushes a price up or down by a tiny amount. A million trades a day. The limit of that process is Brownian motion. Einstein proved this for pollen particles in 1905. The finance world borrowed the math fifty years later. Then three properties that make no sense until you see them derived. Brownian motion crosses zero infinitely often. It never escapes to infinity. And it is nowhere differentiable - with probability one, every path is continuous but has no slope at any point. That last property is why classical calculus breaks completely. Then quadratic variation. For any smooth function, chop an interval into n pieces, square the increments, sum them - the result goes to zero. For Brownian motion it goes to T. The increments are too wild to vanish. That single fact is why Ito's Lemma has a second term that classical calculus does not. Watch the moment he derives it. Taylor expansion applied to a function of Brownian motion. The first term is what you expect. The second term appears precisely because the squared increment does not vanish. Without it, options pricing gives wrong answers. With it, you have Black-Scholes. A quant I know sends this lecture to junior analysts who cannot explain why their pricing model drifts. Says it fixes in ninety minutes what two years of finance courses left open. Free on YouTube, MIT OpenCourseWare, 18.S096. bookmark this and watch later - the math behind every options desk on Wall Street fits on one blackboard, and this is the lecture that shows you why

Lupen

66,176 просмотров • 1 месяц назад

Today we introduce Stochastic Differential Equations (SDEs). I find that the best way to introduce these complex concepts is to look at an application. This is part I of the lecture🙂 We look at the theory behind electromagnetic scattering/radar clutter which leads to anomaly detection on scattering statistics. When a narrowband wave scatters off a messy cloud of particles, the complex field at your receiver is a random phasor sum...at time t you can write the electric field as E_N(t) = Σⱼ₌₁ᴺ e^{iθⱼ(t)}, each term a unit arrow in the complex plane from scatterer j. This is exactly where the magic of Brownian motion appears naturally and in the most reasonable way. Think of all the microscopic chaos...tiny motions, index fluctuations, path jitters, Doppler shifts that shows up as small random kicks to the phases θⱼ(t) over very short times. If you just made θⱼ(t) random in an ad-hoc way (say, resampling independent angles at each time), the field would jump around unrealistically with no temporal structure. Brownian motion is what you get when you let each phase take the continuous-time limit of many tiny, independent kicks...it’s continuous in t, it has the right cumulative variance growth, and it remembers just enough of its past to look physical. So we model each phase as a Brownian walk, θⱼ(t) = θⱼ⁰ + σ_θ Bⱼ(t), with independent Brownian motions Bⱼ(t) and a phase-diffusion rate σ_θ. Brownian motion here isn’t window dressing...it’s the clean way to compress all the small random stuff into a single process that actually matches how the phases wander in time. #StochasticProcesses #BrownianMotion #ItoCalculus #RadarClutter #RayleighScattering #SignalProcessing

Mathelirium

55,618 просмотров • 9 месяцев назад

Today we introduce Stochastic Differential Equations (SDEs), and the main thing to watch for is this: We’ll use Brownian motion as the basic noise source, then see how well-known SDEs drop out of it naturally, without guessing. I still think the best way into these concepts is through an application. We look at the theory behind electromagnetic scattering and radar clutter, which leads straight into anomaly detection on scattering statistics. When a narrowband wave scatters off a messy cloud of particles, the complex field at your receiver is a random phasor sum. At time t you can write the electric field as E_N(t) = Σⱼ₌₁ᴺ e^{iθⱼ(t)}, each term a unit arrow in the complex plane from scatterer j. This is exactly where Brownian motion shows up in the most reasonable way. Think of all the microscopic chaos: tiny motions, index fluctuations, path jitters, Doppler shifts. Over short times, all of that shows up as small random kicks to the phases θⱼ(t). If you made θⱼ(t) random in an ad-hoc way, like resampling a fresh independent angle at every instant, the field would jump around unrealistically with no physical time structure. Brownian motion is what you get when each phase takes the continuous-time limit of many tiny, independent kicks. It’s continuous in t, its variance grows the right way, and it carries just enough temporal structure to look physical. So we model each phase as a Brownian walk, θⱼ(t) = θⱼ⁰ + σ_θ Bⱼ(t), with independent Brownian motions Bⱼ(t) and a phase-diffusion rate σ_θ. Brownian motion here isn’t window dressing. It’s the clean way to compress all the small random stuff into a single process that actually matches how phases wander in time. This is called Rayleigh Scattering, but the same sum of many tiny coherent echoes shows up in lots of places...e.g. wireless multipath fading (phones/Wi-Fi), laser/optical links through atmospheric turbulence, ultrasound speckle in tissue, and sonar/underwater acoustics in rough or bubbly water. #StochasticProcesses #BrownianMotion #ItoCalculus #RadarClutter #RayleighScattering #SignalProcessing

Mathelirium

31,182 просмотров • 8 месяцев назад

Lecture 4 on Calculus of Variations You might wonder...If I’m optimizing a shape...a curve, a surface, a whole path, what does "take the derivative and set it to zero" even mean? Do I take the damn derivative with respect to a curve/surface? 🤔 In normal calculus the variable is a number x, so the reflex is clean...f′(x)=0. In calculus of variations the variable is a whole function...the geometry itself, like a curve y(x) (or a surface z(x,y)). So the derivative can’t be a single slope. It has to be a pointwise sensitivity, i.e. how the objective reacts to tiny local deformations. You’re holding a whole shape, like a curve y(x). Your objective isn’t f(x) anymore, it’s a functional J[y], and as we've seen with our first there examples, usually an integral that depends on the entire curve (often through y and y’). To talk about a “derivative”, you do the only thing that makes sense: you nudge the entire curve by a tiny amount and see how J changes. Pick a wiggle shape η(x). It’s not random...it’s any admissible deformation direction. Admissible just means it obeys the constraints. If the endpoints are fixed, you force η(0)=η(1)=0 so the wiggle doesn’t move the endpoints. Then scale that wiggle by a small number ε and define the perturbed curve yε(x)=y(x)+εη(x). Now treat ε like the usual scalar in a Taylor expansion. As ε→0, J[y+εη] expands as J[y+εη] = J[y] + ε · (first-order term depending linearly on η) + o(ε). So the difference is J[y+εη] - J[y] = ε · (linear functional of η) + o(ε). For the standard integral of a Lagrangian problems, that linear functional can be written as an inner product with some function of x: J[y+εη] - J[y] = ε ∫ (δJ/δy)(x) η(x) dx + o(ε). That’s the definition-level meaning of δJ/δy: it’s the unique pointwise sensitivity function that makes this identity true for every admissible η. If δJ/δy is positive at some x, then choosing η negative there decreases J; if δJ/δy is negative there, pushing y upward locally decreases J. It’s literally a map along the curve saying push this way to go downhill. Now translate “set the derivative to zero.” At a minimizer y*, the first-order change must vanish for every admissible wiggle: J[y*+εη] − J[y*] = o(ε) for all η. Plug in the expansion and the ε-term must be zero: ∫ (δJ/δy)(x) η(x) dx = 0 for all admissible η. Here’s the crucial logic step: the only way an integral against every test function η can be zero is if the integrand itself is zero (in the usual sense used in analysis). So you get δJ/δy = 0. For the common case J[y]=∫ L(x, y, y’) dx, you can compute δJ/δy explicitly and it becomes the Euler–Lagrange expression δJ/δy = ∂L/∂y − d/dx(∂L/∂y’). So if you name the Euler–Lagrange residual as “left-hand side” R(x) = ∂L/∂y − d/dx(∂L/∂y’), then “set the derivative to zero” is exactly R(x)=0. That’s why animation works so well. You don’t have to solve R=0 in one shot. You can evolve the curve in an artificial time τ by pushing it in the downhill direction: ∂y/∂τ = −R(y). Where the residual is large, the curve moves a lot; as the residual drains toward zero, the motion dies out and the curve settles into an extremal. In our animations, we start from an intentionally ugly curve/surface. Frame by frame the functional drops, the residual drains away, and the geometry relaxes into an extremal. #CalculusOfVariations #EulerLagrange #FunctionalDerivative #GradientFlow #Optimization #MathAnimation

Mathelirium

12,186 просмотров • 9 месяцев назад

Ask anyone who’s taken a course in Ordinary Differential Equations (ODEs) what a solution to an ODE represents geometrically, and most of them won’t have a clean answer. When I first took ordinary differential equations, the pattern was always the same. Early on it turns into a speedrun of methods: separation of variables, integrating factors, variation of parameters, Bernoulli, exact equations. Then pretty quickly the course slides into hammer-picking. Spot the form, apply the recipe, move on. Too mechanical! And the real problem is what you don’t walk away with. You leave with a toolkit, but without a feel for what a differential equation even is, especially geometrically. That matters because in real modeling the equations you meet are rarely nice enough to reward memorised recipes. So you get trained to solve toy forms, while the actual subject stays blurry. The behavior. The flow. The shape of solutions. It wasn't until I watched the first lecture of Professor Arthur Mattuck that I realized I didn’t actually know what a solution to a differential equation represents geometrically. His point is almost embarrassingly simple. A first-order ODE is a slope field, and a solution is a curve that stays tangent to that field everywhere. The math breakdown: Write the ODE as dy/dx = f(x,y). At each point (x,y), attach a tiny line segment with slope f(x,y). A function y = y₁(x) is a solution exactly when its graph follows those slopes. At every x, the slope of the curve equals the slope prescribed by the field at the point on the curve. That’s the one line that ties both viewpoints together: y₁′(x) = f(x, y₁(x)). So solving the ODE and drawing an integral curve are the same statement in two languages. Once you see that, you stop obsessing over whether you can write y(x) in closed form. You start asking the questions that actually matter. Where do solutions flow. Where do they get trapped. Where do they blow up. Where does existence or uniqueness fail because the field isn’t even defined? That’s the perspective shift I wish every ODE course forces early. It’s also why I keep pairing math with animation. #DifferentialEquations #ODEs #VectorFields #AppliedMathematics #Mathematics #

Mathelirium

40,841 просмотров • 8 месяцев назад

Do you actually know what convex optimization is in the geometric, guarantee-theoretic sense or have you only met it through solvers and loss curves? Convexity is rare comfort in optimization...there are no spurious local minima, no surprise traps, and inequalities you can use like tools instead of prayers. So, what is this convexity? Let x = (x₁, x₂) and let f(x) be convex. Plot the surface z = f(x). Pick a contact point x₀. The local slope is the gradient p = ∇f(x₀). That p is exactly the data that defines the supporting plane: z = f(x₀) + p · (x − x₀). Thus, f is said to be convex because for every x, f(x) ≥ f(x₀) + p · (x − x₀). So the plane at x₀ can slide under the surface, but it never slices through it. Not near the point...everywhere. Now for here is the interesting part: The slope becomes a coordinate system! Rewrite the same plane as z = p · x − b, where b is the offset. Because the plane passes through (x₀, f(x₀)), the offset is forced to be b = p · x₀ − f(x₀). And that number isn’t just geometry trivia. It’s the convex conjugate: f*(p) = sup over x ( p · x − f(x) ). At a differentiable contact point, the supporting plane touches f tightly enough that the supremum is achieved at x₀, giving the identity f*(p) = p · x₀ − f(x₀) when p = ∇f(x₀). So one moving contact point gives two linked readouts: primal position x₀ dual position (slope) p = ∇f(x₀) dual offset f*(p) One surface. Two worlds. #ConvexOptimization #Optimization #MachineLearning #SignalProcessing #AppliedMath #Engineering

Mathelirium

38,506 просмотров • 9 месяцев назад

When I first took ordinary differential equations, the pattern was always the same. Week 1 turns into a speedrun of methods: separation of variables, integrating factors, variation of parameters, Bernoulli, exact equations… and by Week 2 or 3 the course has quietly degenerated into hammer-picking. Spot the form, apply the recipe, move on. Mechanical! Fuuuuck!😫😫😫😫 The problem is what you don’t walk away with. You leave with a toolkit, but without a feel for what a differential equation even is, especially geometrically. And that’s a big deal, because in real modeling the equations you meet are rarely nice enough to reward memorized recipes. So you end up trained to solve toy forms, while the actual subject...the behavior, the flow, the shape of solutions stays blurry. This is why I’m biased toward the old-timers. Their old-school way of doing things always surprises me:...they’ll spend time on one idea until it sticks, instead of sprinting through a syllabus checklist. One lecture from them and you start noticing a contrast. A lot of modern teaching feels like "finish the content,". You get marched through techniques, but you’re not left with a single thought that keeps bothering you later...the kind of thought that actually pushes you toward research-level curiosity. MIT OpenCourseWare’s Professor Arthur Mattuck did that to me in his very first ODE lecture. One lecture, and your whole relationship with dy/dx = f(x,y) changes. In this segment, Prof. Mattuck is basically saying: A first-order ODE is a slope field, and a solution is a curve that moves everywhere tangent to that field. The math breakdown Write the ODE as dy/dx = f(x,y). At each point (x,y) you attach a tiny line segment with slope f(x,y). A function y = y₁(x) is a solution exactly when its graph follows those slopes:. At every x, the slope of the curve equals the slope prescribed by the field at the point on the curve. That’s the single line that unifies both viewpoints: y₁′(x) = f(x, y₁(x)). So solving the ODE and drawing an integral curve are the same statement in two languages!👌🏻 Once you see that, you can stop obsessing over whether you can write y(x) in closed form. You can start asking the questions that matter: where do solutions flow, where do they get trapped, where do they blow up, and where does existence/uniqueness fail just because the field isn’t even defined? That’s the perspective shift I wish every ODE course forces early and it’s exactly why I keep pairing math with animation. #DifferentialEquations #ODEs #VectorFields #MathAnimation #Mathematics

Mathelirium

53,338 просмотров • 9 месяцев назад

The Trap in Every Mathematics Lecture If you’ve taken a lot of math courses, you start to recognize a pattern. There’s a moment where the lecturer is warming up with the obvious stuff...add matrices entrywise, scale by α, do the row-column product...and you’re thinking, alright… where is this going? Then you relax. You stop resisting. And right there, they slip in one line that changes how you see the whole subject. When Benedict Gross says "matrices represent linear operators,"he’s telling you to stop treating a matrix as a rectangle of numbers and start treating it as an action. A linear operator is a function T: Rⁿ → Rⁿ that respects two rules: T(u+v)=T(u)+T(v) and T(αu)=αT(u). Once you pick a basis, T is completely determined by where it sends the basis vectors e₁,…,eₙ. Put T(e₁),…,T(eₙ) into columns and you get a matrix A. That is what "A represents T" means...A is the coordinate portrait of the transformation. Now the punchline that makes matrix multiplication feel inevitable. If B represents S and A represents T, then doing S first and then T is the composition T∘S. In coordinates that becomes A(Bx)=(AB)x. So multiplying matrices is really composing transformations. That’s why multiplication is usually not commutative: T∘S is generally not the same transformation as S∘T, and the matrices inherit that noncommutativity. This explains half of Linear Algebra because it tells you what the course is really about...functions that move vectors around, not grids of numbers. A matrix is just the written form of that function once you choose coordinates. Then the rules stop feeling random Multiplying matrices means doing one move and then another, an inverse means you can undo the move, eigenvectors are directions that don’t get turned, and changing basis is just describing the same move in a different language. That one idea makes a lot of linear algebra click. #LinearAlgebra #Matrices #GroupTheory #GLn #MathLectures #Mathematics

Mathelirium

66,892 просмотров • 8 месяцев назад

The Trap in Every Mathematics Lecture If you’ve taken enough math courses, you start noticing the same little move. The lecturer warms up with the obvious stuff, add matrices entrywise, scale by α, do the row-column product, and you’re thinking alright, where is this going. Then you relax. You stop resisting. And right there, they drop one line that quietly rewires the whole subject. When Benedict Gross says matrices represent linear operators, he’s telling you to stop treating a matrix as a rectangle of numbers and start treating it as an action. A linear operator is a function T: ℝⁿ → ℝⁿ that respects two rules: T(u+v) = T(u) + T(v) T(αu) = αT(u) Once you pick a basis, T is completely determined by where it sends the basis vectors e₁,…,eₙ. Put T(e₁),…,T(eₙ) into columns and you get a matrix A. That is what A represents T means. A is the coordinate portrait of the transformation. Now the punchline that makes matrix multiplication feel inevitable. If B represents S and A represents T, then doing S first and then T is the composition T∘S. In coordinates that becomes A(Bx) = (AB)x. So multiplying matrices is really composing transformations. That’s why multiplication is usually not commutative. T∘S is generally not the same transformation as S∘T, and the matrices inherit that noncommutativity. This explains half of linear algebra because it tells you what the course is really about: functions that move vectors around, not grids of numbers. A matrix is just the written form of that function once you choose coordinates. After that, the rules stop feeling random. Multiplying matrices means doing one move and then another. An inverse means you can undo the move. Eigenvectors are directions that don’t get turned. Changing basis is just describing the same move in a different language. One idea, and a lot of linear algebra suddenly clicks. #LinearAlgebra #Matrices #LinearMaps #Eigenvectors #ChangeOfBasis #Mathematics

Mathelirium

133,454 просмотров • 7 месяцев назад