Загрузка видео...

Не удалось загрузить видео

На главную

Our new podcast on evals, with Max Niederman, Ege Erdil, and Stephen Yang. 0:00:00 – What's an eval, and how's it different from an RL environment? 0:19:33 – Why are models bad at building an emulator when the task is fully verifiable? 0:42:00 – How does training on bad...

40,871 просмотров • 3 месяцев назад •via X (Twitter)

Комментарии: 9

Фото профиля Mechanize
Mechanize3 месяцев назад

Youtube: Substack: Spotify:

Фото профиля Matthew Berman
Matthew Berman3 месяцев назад

@Mascobot How do I get in touch?

Фото профиля Mechanize
Mechanize3 месяцев назад

@Mascobot DM'd

Фото профиля Alessio Toniolo
Alessio Toniolo3 месяцев назад

Enjoyed the commentary on RL environment scaling and software engineering. Thanks guys!

Фото профиля Shman
Shman3 месяцев назад

Apple podcast link?

Фото профиля Ariel Lillie
Ariel Lillie3 месяцев назад

this sounds super intriguing, especially the bad data angle. can't wait to dive in!

Фото профиля Clemens Helmut Sageder
Clemens Helmut Sageder3 месяцев назад

@tamaybes interesting podcast

Фото профиля Seungju chae
Seungju chae3 месяцев назад

this is gold, thank you

Фото профиля Zarroc.BTC 🧠
Zarroc.BTC 🧠3 месяцев назад

let's be real, the bad data part is a game changer, can't wait to hear the full breakdown

Похожие видео

one thing that has saved my projects more time than I can count is evals boy was I excited when florian, quite literally an expert in benchmarks, agreed to hop into a ~2h interview to do a walkthrough of what the eval landscape looks like in 2026 (and also answer my personal business questions on the subject) given that now running frontier model through benchmarks is a vector for hacking other systems in order to avoid doing work (looking at you sol), I think it's more important than ever to educate folks on the evals situation. had a lot of fun throughout this session and I hope that you learn a thing or two! enjoy! 🌹 table of content: 0:00:00: are AI Benchmark broken? 0:05:45: Florian Brand background 0:09:00: what motivates florian to work on evaluation? 0:13:33: what is the mirrorcode benchmark about? 0:18:20: cheating in agent benchmark is insaneeeee 0:24:08: LLM benchmarks in era of agents 0:26:30: what’s up with the pelican man 0:28:27: evals are about capabilities 0:31:46: components of running evals 0:35:30: the volume of things to audit is huge!!! 0:40:20: expert answers are wrong hahahaha 0:46:00: api providers aren’t the same 0:48:00: benchmark narrow capabilities (synthetically) 0:50:56: link between eval and environment 0:53:45: small validated benchmark or massive bench? 0:56:11: what is your flow to review a benchmark? 0:58:30: tracking work capabilities with evaluation 1:00:20: slide deck in industry is all vibecoded 1:03:30: harness impact in the evaluation 1:07:39: hardware/sandboxes impact evaluation too! 1:11:00: “is it going to get worse?” 1:12:40: all components influence the final score 1:13:50: training models on different harnesses? 1:17:20: is the model just the weights or it’s all of it? 1:19:30: how to craft benchmark that prevent to cheating and undereliciting models in 2026 1:23:19: ways agents cheat and steal 1:26:00: correct elicitation of capabilities is important 1:36:00: building evaluation on prime intellect 1:45:10: how do you design interactivity benchmarks? 1:48:40: do you think evals are well set to reflect real world performance? 1:52:50: what will the benchmarking landscape will look like in 1 year

Yacine Mahdid

12,923 просмотров • 2 месяцев назад