正在加载视频...
视频加载失败
It is very easy to make mistakes when creating evals for your AI product. Shreya Shankar and I run through the most common mistakes in this talk (with memes 🌶️!) . Chapter summaries below: 00:51 Foundation model benchmarks are not the same as your application evals 03:00 Generic Evals... show more
10 条评论

35% discount to our upcoming course: Youtube version of this video:

@sh_reya While it almost is a meme at this point, „look at your data“ is the most important step. Now that I am reading so many benchmark papers, I immediately look whether they done human validation - if they didn’t, there’s always a follow-up paper which does human filtering

@sh_reya Signed up and excited to learn!

Thanks for making this course. But how do you "monitor" the performance of an AI system in production? >Metrics that indicate the correctness of the systems cannot be calculated in prod, as we don't have GT > The signals one can derive from production, such as implicit & explicit feedback, are often weak and assumed to be true. Would love to hear your thoughts.

@sh_reya We discussed this in the video! In prod you also often need to do A/B testing with additional metrics that are more aligned with business outcomes That part is fairly similar to the same paradigm as traditional ML monitoring

@sh_reya @raindrop_ai

@sh_reya Very true. General metrics might be useless or misleading in particular cases.

@sh_reya Impressive, thanks a lot for sharing.

@sh_reya @bytetweets can this be useful?

@sh_reya So many pitfalls to avoid! Thanks for the insightful talk and chapter summaries 🤩
