正在加载视频...
视频加载失败
bro casually explains RL tuning for LLMs and the three critical components: training, inference, and environments. basically any RLVR algorithm such as GRPO comes down to this super simple concept.
102,426 次观看 • 6 个月前 •via X (Twitter)
0 条评论
暂无评论
原始帖子的评论将显示在这里
