
fintex
@_yusufknl • 4,169 subscribers
AI researcher by day nihilist by nature · crypto · models · chaos
Videos

As someone who ships LLM systems in production, this scaling video is the closest thing to a "why 96% of Claude and GPT-5's weights are literally useless" explainer I've ever seen released for free. Everyone thinks trillion-parameter models need every parameter. They don't. A 2019 pruning experiment proved you can delete 96% of a neural net's weights with zero performance loss - meaning most of Claude and GPT-5 is empty scaffolding around a tiny "winning lottery ticket" network doing all the real work. Bookmark this 18-min video and watch tonight. Same lottery ticket math from 2019 MIT research, now the reason every AI lab wastes 90%+ of its Nvidia budget.
fintex788,103 views • 21 days ago

As someone who's been shipping LLMs since the GPT-2 days, this lecture on cross-entropy from a Stanford math grad is the closest thing to an ML PhD qualifying exam I've ever seen released publicly for free. Everyone thinks language models predict the next word. They don't. They compress language. Once you see the math, you can't unsee it. 33 minutes. Bookmark & watch today.
fintex356,334 views • 2 months ago

As someone who's spent 3 years fine-tuning ML models, this lecture on neural networks from a Stanford math grad is the closest thing to a no-bullshit "day one of ML" briefing I've ever seen released publicly for free. Everyone thinks neural nets are magic. They aren't. They're 13,000 dials - a matrix multiplication, a sigmoid squish, done. Once you see the math, "AI" stops feeling mysterious. 18 minutes. Bookmark & watch today.
fintex332,082 views • 2 months ago

In 1948, Claude Shannon invented the math behind every LLM you use today. He tested it by making his wife guess the next letter in a book. A Stanford-trained mathematician just released a 32-minute walkthrough of this exact history - and why the "next-token prediction" story of GPT-5 is actually wrong. Bookmark & watch this weekend. The alternative is a graduate info-theory course + 3 semesters of your life.
fintex288,233 views • 2 months ago

As someone who ships LLM systems in production, this billiards video is the closest thing to a "why every stable LLM behavior is just a lucky island in a sea of chaos" explainer I've ever seen released for free. Everyone thinks GPT-5 always behaves the same way. It doesn't. Every consistent Claude answer lives on a tiny island of stability inside an infinite chaotic plane - the exact "winning zone" this video maps by bouncing balls off polygons. Bookmark this 11-min video and watch tonight. Same emergent stability math from 2007 chaos research, now the reason a single prompt tweak can drop your LLM into total gibberish.
fintex42,313 views • 12 days ago

As someone who's built with modern LLM architectures, this Laplace transform lecture is the closest thing to a "why Mamba works" explainer I've ever seen released for free. Everyone thinks transformers are the only path to modern LLMs. State space models like Mamba and S4 challenge that using math Laplace built in the 1780s. This video shows why. 25 minutes. Bookmark & watch today.
fintex110,467 views • 2 months ago

As someone who trains and fine-tunes LLMs, this double descent video is the closest thing to a "why every ML textbook was wrong about GPT-5" explainer I've ever seen released for free. Everyone thinks bigger models overfit worse. They don't. Past a certain scale the test error dives again - the exact reason OpenAI and Anthropic keep scaling up instead of down. Bookmark & watch this weekend. The 2019 discovery that quietly forced Stanford professors to rewrite their textbooks and reshaped how every frontier LLM gets trained.
fintex68,769 views • 1 month ago

As someone who reads every transformer paper that drops, this Fourier series lecture is the closest thing to a Google FNet explainer I've ever seen released for free. Everyone thinks attention is what makes transformers powerful. Google replaced attention with a 200-year-old Fourier transform. It nearly matched BERT and ran 7x faster. 25 minutes. Bookmark & watch today. Then read the article below - I broke down the 5 pieces of math the AI hype skips.
fintex101,169 views • 2 months ago

As someone who ships LLM systems in production, this sliding puzzle video is the closest thing to a "why sampling harder never finds the answer" explainer I've ever seen released for free. Everyone thinks a hard problem means a huge search space. It doesn't. Klotski has just 25,955 states, yet random search dies in the same pit every time. That's the exact reason OpenAI and Anthropic pay for guided search instead of more samples. Bookmark & watch this weekend. Same graph traversal, from a wooden toy on a coffee table to every reasoning model burning test-time compute today.
fintex57,343 views • 1 month ago

As someone who ships LLM systems in production, this double pendulum video is the closest thing to a "why one changed word breaks Claude and GPT-5" explainer I've ever seen released for free. Everyone thinks LLMs are deterministic machines. They're not. Every prompt to GPT-5 is a chaotic double pendulum - a tiny perturbation in the initial state (one comma, one whitespace) sends the entire response into a completely different attractor. Bookmark this 26-min video and watch tonight. Same butterfly-effect chaos that broke Poincaré's math, now the exact reason your API calls to Claude feel unreliable.
fintex51,856 views • 1 month ago

As someone who trains and fine-tunes LLMs, this backpropagation video is the closest thing to a "the calc 1 rule that trains every LLM on earth" explainer I've ever seen released for free. Everyone thinks training GPT-5 requires exotic math. It doesn't. The entire $500B AI industry runs on the chain rule from freshman calculus - the exact rule Marvin Minsky said would never work. Bookmark this 30-min video and watch tonight. Same math your high school teacher taught you, now updating every weight in Claude, GPT-5, and Llama.
fintex50,749 views • 1 month ago

Schrodinger's equation, Neural ODEs, and every LLM built on Mamba use one operation: raising a matrix to a power. Sounds like nonsense until you see it. A Stanford math grad just released a 27-minute walkthrough starting from a mass on a spring and building up to quantum mechanics. Publicly free. Bookmark & watch this weekend. Same operation, from 1926 to 2025.
fintex43,310 views • 2 months ago

As someone who ships LLM systems in production, this 22-minute Laplace transform video is the closest thing to a "why Mamba solves what transformers can't" explainer I've ever seen released for free. Everyone thinks scaling transformers is the only path forward. State space models beat them on long context using math Laplace built in 1785. This video shows the trick. Bookmark & watch today. The math is older than the United States.
fintex26,102 views • 2 months ago

As someone who builds with LLMs and diffusion models, this 14-minute heat equation video is the closest thing to a "why noise reversal actually works" explainer I've ever seen released for free. Everyone thinks diffusion is a novel AI concept. It's the heat equation from 1822. The forward process spreads noise like heat, and the reverse process is what your model learns. Bookmark & watch this weekend. Same math, from 19th-century physics to 2025 image generation.
fintex15,604 views • 1 month ago

As someone who trains LLMs, this 20-minute gradient descent video is the closest thing to a "why the memorization vs generalization debate matters" explainer I've ever seen released for free. Everyone thinks neural nets learn features. They often just memorize. The video walks through the 2017 paper that shocked the field: shuffle your labels randomly and modern networks will still fit them perfectly. 25 minutes. Bookmark & watch today.
fintex15,922 views • 2 months ago
No more content to load