Loading video...
Video Failed to Load
Jeff Hinton dismissed gradient descent in the 1970s because he thought models would get stuck. He won the Nobel Prize in 2024 partly for being wrong about that. Every time ChatGPT gets better after a training run, it's because 1.2 billion parameters fell through a wormhole in a loss... show more
128,060 views • 11 days ago •via X (Twitter)
0 Comments
No comments available
Comments from the original post will appear here

