正在加载视频...
视频加载失败
New short course on Reinforcement Learning from Human Feedback! RLHF is one of the key techniques that led to the rise of modern LLMs. It is used to align LLMs with human preferences, to make them more honest, helpful and harmless, by (i) learning a reward function that mimics... show more
10 条评论

This is great - it would be interesting to also learn about DPO (Direct Preference Optimization: Your Language Model is Secretly a Reward Model) with which it is possible to skip the RL step for finetuning LLMs:

Mr. Ng. You actually believe yourself in these tech optimist scenario’s? We don’t need AI’s to have ‘understanding’ of friendly behaviour. We need humans to change their behaviour so that we don’t destroy everything we touch of this planet in the wake of our economic ‘progress’

Excited to dive deeper into RLHF-tuning.

Really excited for this new addition to build end to end Gen AI applications via short courses. 🙌🏼 Thank you @DeepLearningAI @googlecloud 🤗

new rlhf course? awesome! wondering how this can simplify llm tuning for startups. anyone planning to integrate it into their ai strategy?

This is great. I will watch it this weekend.

🙏 Thank you for promoting high quality content!

Super insightful, @AndrewYNg

marvelous

Excellent!
