正在加载视频...

视频加载失败

Check out this week’s top Web3 Gaming NEWS!🔥 0:19 Avalanche announced a major partnership to enhance games that will integrate AI agents. 0:42 OLAGG draws $500 USDC. 1:00 A medieval themed event took place in Brazil. 1:22 MMORPG game launches its creator program. 1:45 First shooter on the Soneium...

10,954 次观看 • 1 年前 •via X (Twitter)

2 条评论

Predzᴸ¹🔺 的头像
Predzᴸ¹🔺1 年前

Featured on this video: @avax @LFG_GTFO @OLAGuildGames @playSIPHER @RavenQuestGame @TempestGuild @WayfindersGG @spellbornegame @PlayOverTrip @soneium @ceden_network @bballverse_gg @AbstractChain @play_ember @playgigaverse @ForgotPlayland

Yandel Krkič 的头像
Yandel Krkič1 年前

What

相关视频

one thing that has saved my projects more time than I can count is evals boy was I excited when florian, quite literally an expert in benchmarks, agreed to hop into a ~2h interview to do a walkthrough of what the eval landscape looks like in 2026 (and also answer my personal business questions on the subject) given that now running frontier model through benchmarks is a vector for hacking other systems in order to avoid doing work (looking at you sol), I think it's more important than ever to educate folks on the evals situation. had a lot of fun throughout this session and I hope that you learn a thing or two! enjoy! 🌹 table of content: 0:00:00: are AI Benchmark broken? 0:05:45: Florian Brand background 0:09:00: what motivates florian to work on evaluation? 0:13:33: what is the mirrorcode benchmark about? 0:18:20: cheating in agent benchmark is insaneeeee 0:24:08: LLM benchmarks in era of agents 0:26:30: what’s up with the pelican man 0:28:27: evals are about capabilities 0:31:46: components of running evals 0:35:30: the volume of things to audit is huge!!! 0:40:20: expert answers are wrong hahahaha 0:46:00: api providers aren’t the same 0:48:00: benchmark narrow capabilities (synthetically) 0:50:56: link between eval and environment 0:53:45: small validated benchmark or massive bench? 0:56:11: what is your flow to review a benchmark? 0:58:30: tracking work capabilities with evaluation 1:00:20: slide deck in industry is all vibecoded 1:03:30: harness impact in the evaluation 1:07:39: hardware/sandboxes impact evaluation too! 1:11:00: “is it going to get worse?” 1:12:40: all components influence the final score 1:13:50: training models on different harnesses? 1:17:20: is the model just the weights or it’s all of it? 1:19:30: how to craft benchmark that prevent to cheating and undereliciting models in 2026 1:23:19: ways agents cheat and steal 1:26:00: correct elicitation of capabilities is important 1:36:00: building evaluation on prime intellect 1:45:10: how do you design interactivity benchmarks? 1:48:40: do you think evals are well set to reflect real world performance? 1:52:50: what will the benchmarking landscape will look like in 1 year

Yacine Mahdid

12,923 次观看 • 1 个月前