正在加载视频...
视频加载失败
Don't waste 2 years learning how Claude and ChatGPT actually work. Stanford just dropped a 1-hour course on the exact pipeline behind them. 0:00 - policy gradient basics 23:02 - PPO for LLM training 55:28 - how models learn chain of thought 1:02:50 - the architecture, explained look at... show more
21,564 次观看 • 17 天前 •via X (Twitter)
0 条评论
暂无评论
原始帖子的评论将显示在这里
