正在加载视频...

视频加载失败

Taleb: "standard deviation is not how much something moves on average." ask anyone - even statisticians, even government agencies - and that's exactly the wrong definition you'll get back. what they're describing is mean absolute deviation. real standard deviation squares the moves first - which quietly hands almost all...

412,623 次观看 • 1 个月前 •via X (Twitter)

0 条评论

暂无评论

原始帖子的评论将显示在这里

相关视频

No author has shaped my worldview more than Nassim Taleb (Nassim Nicholas Taleb). I stumbled on The Black Swan eight years ago, then inhaled the rest of the Incerto. I've increasingly come to appreciate the importance of its ideas to prolonging the human story. Was an honour to speak with Nassim on the podcast. Hard to summarise our conversation, but the timestamps below capture the gist. Enjoy! Timestamps: (0:00:55) - Heuristics for knowing when you're in Mediocristan versus Extremistan. (0:06:06) - Are certain tail exponents intrinsic? (0:10:30) - Why hasn't Universa's tail hedging strategy now been fully priced in? (0:11:52) - Does the power law distribution of startup returns mean VCs should concentrate their bets, or spray and pray? (0:15:20) - Nassim's 30-minute take on the field of behavioural economics. (0:48:57) - Nassim's 20-minute take on superforecasting. (1:11:03) - The Precautionary Principle and AI. (1:17:28) - What are LLMs doing? (1:23:10) - War, violence, & "the empirical mean is not the real mean". (1:39:17) - Covid, & how Western governments think about tail risk. (1:43:38) - What's the most important thing people in social science get wrong about correlation? (1:52:58) - How does Nassim explain the perspicacity of the Russian school of probability? (1:55:56) - Why doesn't Hayek's knowledge argument extend to prediction markets? (1:58:26) - If mean absolute deviation is a better measure than standard deviation, why has the latter become commonplace? (2:01:09) - Nassim's next book, and what he's up to at the moment.

Joseph Noel Walker

577,821 次观看 • 1 年前

Variational Autoencoder by hand ✍️ ~ 11 steps walkthrough below A VAE learns the structure of your data, the mean and variance of its hidden features, and then generates new data from that structure. A GAN only learns to fool a discriminator. It can make convincing fakes without ever knowing what the data is really made of. That is the difference, and it is the whole reason VAEs matter. In 2024 ICLR gave its first ever Test of Time Award to the VAE paper, "Auto-Encoding Variational Bayes" by Diederik Kingma and Max Welling, ten years on. How does it work? Goal: encode three inputs into a distribution, sample from it, decode it back, and read every loss gradient off the page. = 1. Given = Three training examples X1, X2, X3, copied to the bottom as their own targets. Reconstructing your own input is what puts the "auto", meaning self, in autoencoder. = 2. Encoder, layer 1 = Let us multiply the inputs by weights and biases, then apply ReLU, crossing out every negative. = 3. Mean and standard deviation = We multiply the features by two more weight sets. The first predicts the means μ of the latent distributions, the second their standard deviations σ. = 4. A random offset = Let us sample ε from a standard normal, mean 0 and variance 1, and multiply it by σ. This is a random step away from the mean, scaled by how uncertain each feature is. = 5. Mean plus offset = We add the offset back onto μ, and these become the decoder's inputs. Keeping the randomness out in ε is the reparameterization trick: it lets gradients flow straight through the sampling. = 6. Decoder, layer 1 = Let us multiply by weights and biases and apply ReLU again. Here -4 is crossed out. = 7. Decoder, layer 2 = We multiply once more. The output Y is the decoder's attempt to rebuild X from the sampled distribution. = 8. Gradient for the mean = Let us push μ toward 0. A lot of math, the SGVB estimator, collapses the KL gradient to simply μ itself. = 9. Gradient for the standard deviation = We want σ to approach 1. = 10. And its formula = That same math simplifies the gradient to σ minus 1/σ. = 11. Reconstruction gradient = We want the reconstruction Y to match the input X. Mean squared error simplifies its gradient to Y minus X. Takeaway: the two gradients you just calculated each sit at the heart of a modern method, so one VAE teaches you both. The KL divergence is the penalty RLHF like GRPO uses to keep a fine-tuned model from drifting off its base. The reconstruction loss, plain mean squared error, is exactly what trains a diffusion model to denoise. Draw one VAE by hand and you have quietly learned the core of both. 💾 Save this post!

Tom Yeh

17,011 次观看 • 19 天前

Marc Andreessen says raw intelligence might be the worst qualification for leadership — and it changes everything about how we should think about AI. "If the leader is more than one standard deviation of IQ away from the followers, it's a real problem." Andreessen points to the US military, one of the earliest and most rigorous adopters of IQ testing, as the source of this insight. They slot people into specialties and leadership roles based on IQ scores. And over the years, they kept seeing the same pattern. A leader who is significantly less intelligent than their people struggles to model how those people think. That part is intuitive. But the reverse turns out to be equally true. "It's actually very hard for very smart people to model the internal thought processes of even moderately smart people." A leader who is two standard deviations above the norm of the organisation they're running also loses theory of mind, that ability to hold an accurate model of what's happening inside someone else's head. The gap is too wide in both directions. Andreessen then takes this to its logical conclusion: "If you had a person or a machine that had a thousand IQ or something like it, its understanding of reality would be so alien to the people or the things that it was managing that it wouldn't even be able to connect in any sort of realistic way." An AI that vastly outthinks every human in the room isn't positioned to lead those humans. It's positioned to be completely incomprehensible to them. Leadership has never really been an intelligence problem. It's a connection problem. And no amount of raw intelligence closes that gap — past a certain point, it only widens it. The world will not be run by the smartest thing in the room for a long time. Maybe ever.

Big Brain AI

366,029 次观看 • 4 个月前

This guy is making +1% on EVERY TRADE Sounds small, I know. But he already made $50k profit (started with $200 only) How? He repeats this +1% trade 30 times per day. Almost zero risk. And I found a way to make 100x more trading same markets. This is either genius or it should be illegal. But this trader has been quietly running the cleanest strategy on Polymarket for months. Just one repeatable process: > Find markets where the outcome is already decided > Buy shares at 95-99¢ > Collect the gap to $1.00 at resolution > Wake up tomorrow and do it again His PnL curve looks like a savings account that somehow prints 33% daily. Wallet: 5% per trade sounds boring until you run the parlay math. Here's exactly what changes with PolyParlay: Standard bond trading on $1,000: > 4 markets at 96¢ traded separately > +4% collected four times > Total: $1,060 (you made $60, congrats) Same $1,000 combined into one parlay: > 0.96 × 0.96 × 0.96 × 0.96 = boosted PnL > 1.27x payout multiplier > Total: $1,270 in one single position Now add 2 mispriced entries at 70-80¢ into that same parlay: > Payout multiplier jumps to 3-10x > Same $1,000 becomes $3,000-$10,000 after all hit That's the gap between grinding $160 daily and waking up to $3,000+ on the same capital. Bond markets eliminate the risk. Parlay structure eliminates the ceiling. Both together is what $50k months actually look like. Parlay bot link: This is the only bot for Polymarket parlay trading. Zero stress. Daily yield. Exponential upside.

Oracle Boar

13,926 次观看 • 3 个月前