Video wird geladen...
Video konnte nicht geladen werden
Don't waste 2 years learning how Claude and ChatGPT actually work. Stanford just dropped a 1-hour course on the exact pipeline behind them. 0:00 - policy gradient basics 23:02 - PPO for LLM training 55:28 - how models learn chain of thought 1:02:50 - the architecture, explained look at... show more
21,564 Aufrufe • vor 18 Tagen •via X (Twitter)
0 Kommentare
Keine Kommentare verfügbar
Kommentare vom Original-Post werden hier angezeigt
