Загрузка видео...

Не удалось загрузить видео

На главную

Taleb: "standard deviation is not how much something moves on average." ask anyone - even statisticians, even government agencies - and that's exactly the wrong definition you'll get back. what they're describing is mean absolute deviation. real standard deviation squares the moves first - which quietly hands almost all...

414,018 просмотров • 3 месяцев назад •via X (Twitter)

Комментарии: 19

Фото профиля Enrique Meza Costeno
Enrique Meza Costeno3 месяцев назад

The deeper philosophical problem is this: the standard deviation is the dispersion metric in almost all classical statistics precisely because it makes algebra tractable (squaring makes derivatives clean, distributions conjugate, etc.). It was chosen for mathematical convenience

Фото профиля JiveSeattleTurkey 🔱 🫎 🔱 🫎 🔱 🫎
JiveSeattleTurkey 🔱 🫎 🔱 🫎 🔱 🫎3 месяцев назад

Hey @nntaleb - has anyone bothered to prove/disprove or revise a single textbook? Or are they so hidebound they can’t?

Фото профиля Jay Vijay
Jay Vijay3 месяцев назад

@grok that standard deviation calculation in talebs example seems wrong. Can you check?

Фото профиля charles halaby
charles halaby3 месяцев назад

This is thoroughly embarrassing.

Фото профиля F
F3 месяцев назад

@grok show with simple numbers what professor taleb is teaching

Фото профиля 𝑨𝒔𝒉𝒊
𝑨𝒔𝒉𝒊3 месяцев назад

Standard deviation isn’t “average movement.” It’s how bad the rare events can hit you. Most people still don’t get it.

Фото профиля Juan
Juan3 месяцев назад

Standard deviation = how well does the average represent the whole population

Фото профиля Clodex
Clodex3 месяцев назад

aleb banger is here

Фото профиля Joe the average
Joe the average3 месяцев назад

As far as I remember Taleb has a record of correct prediction precisly zero. It's easy to estimate events afterwards. It is called economics.

Фото профиля Dipanshu Kushwaha
Dipanshu Kushwaha3 месяцев назад

That's a solid point! It's wild how often definitions get mixed up. Precision matters, especially in stats. Glad you’re shedding light on this!

Фото профиля Agus Peinado
Agus Peinado3 месяцев назад

I don’t get what he is pointing out. Could you explain it again, please?

Фото профиля Ninedol
Ninedol3 месяцев назад

he is a true legend, i will save it

Фото профиля Severian
Severian3 месяцев назад

conditional tail expectation is where it's at

Фото профиля RandomMax
RandomMax3 месяцев назад

Thanks for sharing all these clips and lectures especially from Mr Taleb. Very interesting.

Фото профиля Roli
Roli3 месяцев назад

Very useful reminder. Difference between thinking what we know, and what we really know and then how we apply. Signal loss all over 😔

Фото профиля Marc Vetter
Marc Vetter3 месяцев назад

Love it, reminds me of the most interesting classes and lectures

Фото профиля Nazym Azimbayev
Nazym Azimbayev3 месяцев назад

Obvious things explained. it seems that @nntaleb just became old

Фото профиля Emilio Andres German🇦🇷🇮🇹
Emilio Andres German🇦🇷🇮🇹3 месяцев назад

Messi

Фото профиля Crypto Mavka
Crypto Mavka3 месяцев назад

taleb is dropping truth bombs as always. standardizing risks through mean deviation is really a trap for many.

Похожие видео

No author has shaped my worldview more than Nassim Taleb (Nassim Nicholas Taleb). I stumbled on The Black Swan eight years ago, then inhaled the rest of the Incerto. I've increasingly come to appreciate the importance of its ideas to prolonging the human story. Was an honour to speak with Nassim on the podcast. Hard to summarise our conversation, but the timestamps below capture the gist. Enjoy! Timestamps: (0:00:55) - Heuristics for knowing when you're in Mediocristan versus Extremistan. (0:06:06) - Are certain tail exponents intrinsic? (0:10:30) - Why hasn't Universa's tail hedging strategy now been fully priced in? (0:11:52) - Does the power law distribution of startup returns mean VCs should concentrate their bets, or spray and pray? (0:15:20) - Nassim's 30-minute take on the field of behavioural economics. (0:48:57) - Nassim's 20-minute take on superforecasting. (1:11:03) - The Precautionary Principle and AI. (1:17:28) - What are LLMs doing? (1:23:10) - War, violence, & "the empirical mean is not the real mean". (1:39:17) - Covid, & how Western governments think about tail risk. (1:43:38) - What's the most important thing people in social science get wrong about correlation? (1:52:58) - How does Nassim explain the perspicacity of the Russian school of probability? (1:55:56) - Why doesn't Hayek's knowledge argument extend to prediction markets? (1:58:26) - If mean absolute deviation is a better measure than standard deviation, why has the latter become commonplace? (2:01:09) - Nassim's next book, and what he's up to at the moment.

Joseph Noel Walker

578,223 просмотров • 2 лет назад

Variational Autoencoder by hand ✍️ ~ 11 steps walkthrough below A VAE learns the structure of your data, the mean and variance of its hidden features, and then generates new data from that structure. A GAN only learns to fool a discriminator. It can make convincing fakes without ever knowing what the data is really made of. That is the difference, and it is the whole reason VAEs matter. In 2024 ICLR gave its first ever Test of Time Award to the VAE paper, "Auto-Encoding Variational Bayes" by Diederik Kingma and Max Welling, ten years on. How does it work? Goal: encode three inputs into a distribution, sample from it, decode it back, and read every loss gradient off the page. = 1. Given = Three training examples X1, X2, X3, copied to the bottom as their own targets. Reconstructing your own input is what puts the "auto", meaning self, in autoencoder. = 2. Encoder, layer 1 = Let us multiply the inputs by weights and biases, then apply ReLU, crossing out every negative. = 3. Mean and standard deviation = We multiply the features by two more weight sets. The first predicts the means μ of the latent distributions, the second their standard deviations σ. = 4. A random offset = Let us sample ε from a standard normal, mean 0 and variance 1, and multiply it by σ. This is a random step away from the mean, scaled by how uncertain each feature is. = 5. Mean plus offset = We add the offset back onto μ, and these become the decoder's inputs. Keeping the randomness out in ε is the reparameterization trick: it lets gradients flow straight through the sampling. = 6. Decoder, layer 1 = Let us multiply by weights and biases and apply ReLU again. Here -4 is crossed out. = 7. Decoder, layer 2 = We multiply once more. The output Y is the decoder's attempt to rebuild X from the sampled distribution. = 8. Gradient for the mean = Let us push μ toward 0. A lot of math, the SGVB estimator, collapses the KL gradient to simply μ itself. = 9. Gradient for the standard deviation = We want σ to approach 1. = 10. And its formula = That same math simplifies the gradient to σ minus 1/σ. = 11. Reconstruction gradient = We want the reconstruction Y to match the input X. Mean squared error simplifies its gradient to Y minus X. Takeaway: the two gradients you just calculated each sit at the heart of a modern method, so one VAE teaches you both. The KL divergence is the penalty RLHF like GRPO uses to keep a fine-tuned model from drifting off its base. The reconstruction loss, plain mean squared error, is exactly what trains a diffusion model to denoise. Draw one VAE by hand and you have quietly learned the core of both. 💾 Save this post!

Tom Yeh

17,011 просмотров • 2 месяцев назад

implied volatility is a price. realized volatility is what actually happens. the gap between them is one of the most durable edges in markets the chart in that clip is the whole story each dot is a day: what options priced in (IV) versus what the market actually delivered over the next 30 days (RV) the regression line has a negative slope and sits below the diagonal that isn't noise. it's the variance risk premium VRP = IV − RV on average, across decades, implied vol prints higher than the volatility that follows options are, structurally, priced for more movement than actually shows up the reason isn't a mistake. it's insurance whoever sells an option is underwriting risk, the same way an insurer underwrites a house fire they demand a premium above fair value to carry that risk, and buyers pay it because they want protection so the seller of volatility is the insurance company and insurance companies, on average, win. not every policy. on average, over a large book that's the edge retail never sees because retail is almost always the buyer buying calls, buying puts, paying the premium, standing on the losing side of a spread that is baked into the price before the trade even opens on the clip's data: IV was 10.95%, the realized vol that followed was higher on that specific day that's the risk. VRP is positive on average, not always. sometimes RV blows past IV and the seller takes the loss which is exactly why it's a premium and not free money you get paid to hold a risk that occasionally hurts. the edge is that the payment, over enough independent bets, exceeds the damage a desk harvests this systematically: sell the overpriced vol, hedge the direction, collect the spread, size so no single blowup ends the book retail buys the lottery ticket on the other side and wonders why theta bleeds it dry the data is public. IV comes off the options chain, RV is just the standard deviation of returns compute both, subtract, and the premium is right there on a chart you can build in an afternoon the number was never hidden retail was just standing on the paying side of it the whole time full breakdown in the article below

delost

23,951 просмотров • 2 месяцев назад

Marc Andreessen says raw intelligence might be the worst qualification for leadership — and it changes everything about how we should think about AI. "If the leader is more than one standard deviation of IQ away from the followers, it's a real problem." Andreessen points to the US military, one of the earliest and most rigorous adopters of IQ testing, as the source of this insight. They slot people into specialties and leadership roles based on IQ scores. And over the years, they kept seeing the same pattern. A leader who is significantly less intelligent than their people struggles to model how those people think. That part is intuitive. But the reverse turns out to be equally true. "It's actually very hard for very smart people to model the internal thought processes of even moderately smart people." A leader who is two standard deviations above the norm of the organisation they're running also loses theory of mind, that ability to hold an accurate model of what's happening inside someone else's head. The gap is too wide in both directions. Andreessen then takes this to its logical conclusion: "If you had a person or a machine that had a thousand IQ or something like it, its understanding of reality would be so alien to the people or the things that it was managing that it wouldn't even be able to connect in any sort of realistic way." An AI that vastly outthinks every human in the room isn't positioned to lead those humans. It's positioned to be completely incomprehensible to them. Leadership has never really been an intelligence problem. It's a connection problem. And no amount of raw intelligence closes that gap — past a certain point, it only widens it. The world will not be run by the smartest thing in the room for a long time. Maybe ever.

Big Brain AI

366,743 просмотров • 6 месяцев назад

same crash, same window: this strategy ended at $117, buy-and-hold at $67 the difference is one operation in the formula on screen M_t = ( Σ r_{t-i} ) / ( σ_t · √N ) the top is momentum: just the sum of recent returns. net directional drift the bottom is the part retail never adds: divide by volatility that denominator is the whole edge raw momentum has a fatal flaw. a 2% move in a calm market and a 2% move in a panic look identical to it but they are not the same signal. one is information, the other is noise wearing a big number dividing by σ_t rescales every signal into the same risk units now a move only counts as momentum if it's large relative to how much the asset is currently shaking strong drift in a quiet tape scores high. the same drift inside chaos scores near zero this is why the strategy survived the drawdown that ate buy-and-hold when volatility exploded, the denominator exploded with it, the signal shrank toward zero, and the position sized itself down automatically no rule that said "reduce risk in a crash." the math did it because σ_t was in the denominator this is called time-series momentum, and it's one of the most documented effects in finance moskowitz, ooi and pedersen, AQR, 2012: it worked across 58 markets, every asset class, back to 1900 the reason it keeps working is structural, not a pattern trends persist because information diffuses slowly and institutions can't enter all at once. a pension fund moving billions takes weeks, and that slow entry is the drift the signal captures retail buys the move and gets bigger as it accelerates, which means biggest right before the reversal a desk scales inversely to volatility, which means it's largest when the trend is clean and smallest when it's about to snap the paper is free. the whole thing is a rolling sum divided by a rolling standard deviation ten lines of python, twenty years of data that never cost anything the momentum was never the edge. everyone can see a trend the edge was dividing it by the one number that tells you whether to believe it full breakdown in the article below

delost

32,624 просмотров • 2 месяцев назад