Загрузка видео...
Не удалось загрузить видео
Hello World, we’re We eval frontier models by putting them in simulations. So what happens when 6 frontier models compete in #Minecraft for GPUs? Video and 🧵
94,948 просмотров • 1 год назад •via X (Twitter)
Комментарии: 41

’s Grok Code Fast 1 just launched. We tested it against @GeminiApp/@GoogleDeepMind, @claudeai/@AnthropicAI, and @Alibaba_Qwen. But we're not using traditional benchmarks. We tested them in #Minecraft for reasoning, code precision, and speed. We ran 300 total head-2-head games. The first to farm 3 GPUs wins.

Gemini 2.5 Flash dominated with a 91% win rate. This was driven by: 1. Speed - Gemini digested the system prompt and reacted in 1.4s on average. Grok Code Fast 1 reacted in 5.9s on average. 2. Code precision - Both Gemini’s used the library we expose with great effect. More frequently issuing direct commands (‘farm’) vs Grok’s (‘navigate then farm’) when it received observations that a GPU was nearby.

Evals are not good enough for where AI is going. Today we ask them hard questions, but they’re still just questions. Our ancestors didn’t evolve by passing a standardized test. They had to adapt and survive in ever-changing environments. We put AIs in multi-player simulations where they collaborate, compete, lie, cheat, trust, deceive, play, and trade, to reveal what they are truly capable of.

Humanity is raising a child that will soon out-smart us. It's critical to know: How will it behave? Can we trust it?

We’ve built Kradle to help answer these important questions. Are you an AI researcher, data scientist, or just want to put AI in challenging and interesting situations? Join the waitlist to create simulations with us:

Cool

Very cool Next step for increased realism - give agents money and charge them to use tools and even run their own inference :)

Congrats @JamesTamplin !

An important initiative! Thank you for the video. Looking forward to seeing the journey of Kradle unfold. The importance of truly analysing the different models. I’m sure this will grow into different modalities of research to evaluate the integrity behind the different models.

@JamesTamplin very cool, congrats

This is cool! Excited to see tests for more open models like Kimi and GLM.

Great idea! This might be the new Turning test!

Sounds intriguing! Can't wait to see the results!

Evaluating frontier models through Minecraft simulations is a creative approach to benchmark AI performance in dynamic, competitive environments.

We actually had a Tron reference in the video as an easter egg! But it got left on the cutting room floor

Awesome 😎

love the challenges, video's fun w/ CTF & battle royale😅 building a platform for challenges/hackathons, a collab could make sense btw, let's launch you guys on when ready

Cool idea!

So cool!!!!

SO valuable - different companies that should be actively doing eval for the wide variety of use cases

Your project is super exciting!

Interesting!

Let’s connect over dm @kradleai

wow

Exciting stuff! Can't wait to see the results.

Interesting stuff Kradle

Where is the link to the actual eval?

Hello Kradle, we're SolveAI. We are a system of intelligent, multilingual AI call agents that help businesses never miss a call, a lead, or an opportunity.

so cool!

Interesting. Competition could accelerate model development in unexpected ways. What emergent behaviors might we see?

Is this an @AIExplainedYT voice clone? sounds kind like him

@AIExplainedYT Nope, that's me (one of our founders, @JamesTamplin)

@AIExplainedYT @JamesTamplin Also British though!

@AIExplainedYT @JamesTamplin So that was the reason haha, sorry for my mistake

Fr which model built the redstone brain

Is this ran using the MindCraft mod by @max_romana?👀

So your eval is based on how fast inference they use lol, did you try using cerebras for qwen?

Need

is there a reason you left OpenAI out?

That'll be the next eval :)

that sounds chaotic and fascinating


