Загрузка видео...

Не удалось загрузить видео

На главную

Mental map of Markov Chain Monte Carlo (MCMC) algorithms, and analogous machine learning (ML) algorithms [dashed = especially loose analogy]. Grey boxes are basic tools, and each arrow is annotated with the "delta" between algorithms.

100,691 просмотров • 1 год назад •via X (Twitter)

Комментарии: 10

Фото профиля Keenan Crane
Keenan Crane1 год назад

I'm developing this "map" as part of a new course on Monte Carlo Methods and Applications, taught at @CarnegieMellon this fall: (See also last year's offering:

Фото профиля Keenan Crane
Keenan Crane1 год назад

Very glad to hear from seasoned travelers about errors in my cartography! Likewise glad to know about papers that make some of the MCMC <=> ML connections precise/rigorous. I know a bunch that touch on these connections, like Welling & Teh 2011:

Фото профиля Keenan Crane
Keenan Crane1 год назад

(For related posts, see the tag #MCMA2023.)

Фото профиля Kyle Cranmer
Kyle Cranmer1 год назад

Recently, there are MCMC (and importance sampling) algorithms that use neural density estimation to act as a good proposal distribution.

Фото профиля Artur Chakhvadze
Artur Chakhvadze1 год назад

Is there a good derivation of Langevin Monte-Carlo from first principles which is independent from Hamiltonian Dynamics? The only one I know is “We just do one-step HMC and by making step sizes smaller over time we make sure that acceptance probability approaches 1”

Фото профиля Keenan Crane
Keenan Crane1 год назад

Yeah, definitely. You don't need the symplectic picture do develop LMC: just "gradient descent plus noise" (more or less). If you don't want it to totally fall out of the sky, you can use motivation from physics. The tough part is proving it *works*, which takes 1/2 a semester!

Фото профиля Amit Patel
Amit Patel1 год назад

I'd love a non-animated version of this so I can read it before the animation resets ;)

Фото профиля Keenan Crane
Keenan Crane1 год назад

I think if you fullscreen it you may get a pause button. (But I totally hear you!)

Фото профиля Bassel Mabsout
Bassel Mabsout1 год назад

I can see how the cooling schedule captures the essence of simulated annealing, but it's supposed to be a gradientless metaheuristic, is there some more general definition of simulated annealing that can be considered the global optimizer in conjunction with some local optimizer?

Фото профиля Keenan Crane
Keenan Crane1 год назад

Yeah, that’s actually a good point. That box should probably be something like Simulated Annealing with Langevin Sampling. It’s indeed important that simulated annealing (e.g., with Metropolis/normal proposals) can be applied without derivatives.

Похожие видео

25 algorithms every programmer should know: Let's start with my top favorite 10. If nothing else, you should read about these algorithms and have a good idea of how they work: 1. Linear search to find an element in a list 2. Binary search to find an element on a sorted list 3. Bubble sort to sort a list 4. Merge sort will also sort lists 5. Quicksort to sort the list and do it fast 6. Dijkstra to find the shortest path in a graph 7. Breadth-first Search (BFS) for trees or graphs 8. Depth-first search (DFS) for trees or graphs 9. Huffman for doing data compression 10. Anything related to dynamic programming Learning about algorithms is like getting tattoos: you never have enough. Here are another 5 algorithms that will help you go beyond the basics: 11. Kruskal for the finding minimum spanning tree 12. Floyd Warshall, shortest paths in a graph 13. Union Find to detect cycles in a graph 14. Bellman-Ford, shortest path in a graph 15. Lee for finding the shortest path in a maze If you are serious about this topic, I recommend learning about algorithms' space and time complexity. People usually refer to this topic as "Big O" notation. You should build a good intuition about the performance of different algorithms and learn how to evaluate them. Machine Learning will rule the next 50 years, so the next 10 algorithms you can't ignore are the following: 16. Linear Regression 17. Logistic Regression 18. Decision Trees 19. Bayes' theorem 20. k-Nearest Neighbors (kNN) 21. Every algorithm related to neural networks 22. K-means 23. Random forest 24. Gradient boosting algorithms 25. Any dimensionality reduction algorithm (PCA, for instance) There are many more mind-blowing algorithms! I haven't found a better way to understand how computers work from a first-principles point of view than reading about different algorithms. Take a look at the attached video.

Santiago

274,092 просмотров • 2 лет назад

The most important tool in Probability and Statistics - Markov Chain Monte Carlo (MCMC) Method Fresh out of undergraduate Probability and Stats courses, it’s easy to feel invincible. You’ve tamed Gaussians, gammas, betas, all those neat closed-form toy distributions. Then research hits and you meet the harsher truth. Real posteriors and energy landscapes are jagged, asymmetric, multimodal, and too high-dimensional to integrate or sample from directly. You can’t compute the normalising constant. You can’t do the integrals by hand. And i.i.d. samples are basically science fiction. Markov Chain Monte Carlo is the hack we invented to survive that reality. Instead of drawing perfect samples, you send a carefully designed random walk wandering through the landscape, then use its long-run positions as your window into the target distribution. Here’s the problem. Standard trace plots and diagnostics can still cheerfully lie to you. High-dimensional geometry can make a chain that looks healthy while it’s effectively frozen. Multimodal targets, bad tuning, and hidden correlations can quietly wreck your posterior summaries. This series is about those blind spots. We’ll use visuals like this one to show how MCMC actually moves, where the guarantees get slippery, and how to think clearly about convergence and diagnostics in serious Bayesian, physics, and ML work. #BayesianInference #MCMC #MonteCarloMethods #ProbabilityLandscape #StatisticsEducation #ComputationalScience

Mathelirium

32,296 просмотров • 6 месяцев назад

Everybody is talking about recursive self-improvement (RSI) and meta learning. Here is my old 2020 talk about this [1]. It has aged well. Example: humans still define the starts & ends of trials of many modern meta learners. My RSI systems since 1994 LEARN to (re)define them [2]! [1] Meta Learning Machines in a Single Lifelong Trial (talk for workshops at ICML 2020 and NeurIPS 2021, based on earlier talks since 1994). Abstract: the most widely used machine learning algorithms were designed by humans and thus are hindered by our cognitive biases and limitations. Can we also construct meta learning algorithms that can learn better learning algorithms so that our self-improving AIs have no limits other than those inherited from computability and physics? This question has been a main driver of my research since I wrote a thesis on it in 1987 [2]. Here I summarize our work on meta reinforcement learning with self-modifying policies in a single lifelong trial (since 1994), and mathematically optimal meta-learning through the self-referential Gödel Machine (since 2003). Many additional publications on meta-learning since 1987 can be found in the RSI overview [2]. [2] J. Schmidhuber (AI Blog, 2020-2025). 1/3 century anniversary of first publication on recursive self-improvement (RSI) and meta learning machines that learn to learn (1987). For its cover I drew a robot that bootstraps itself. 1992-: gradient descent-based neural meta learning. 1994-: meta reinforcement learning with self-modifying policies. 1997: meta RL plus artificial curiosity and intrinsic motivation. 2002-: asymptotically optimal meta learning for curriculum learning. 2003-: mathematically optimal Gödel Machine. 2020-: new stuff!

Jürgen Schmidhuber

232,164 просмотров • 5 месяцев назад