正在加载视频...
视频加载失败
RLHF by hand ✍️ ~ 15 steps walkthrough below Train a model on human text and it inherits human bias. It will assume a doctor is a "him", because the data says so. RLHF is the correction. A human marks one preference, doc is them over doc is him,... show more
14,128 次观看 • 6 天前 •via X (Twitter)
0 条评论
暂无评论
原始帖子的评论将显示在这里
