Loading video...

Video Failed to Load

Go Home

LLaMA-4 Maverick performs well on reasoning benchmarks and ranks 2nd on the Chatbot Arena, yet its true performance remains controversial. What if we put them in a transparent gaming environment? 🎮 Our benchmark tells a different story...🤔 Will true intelligence shine through play? Let’s find out 👇

39,141 views • 1 year ago •via X (Twitter)

4 Comments

Hao AI Lab's profile picture
Hao AI Lab1 year ago

We directly compared LLaMA-4 Maverick with its reported peers. Despite of LLaMA-4’s strong performance on static benchmarks, in 2048, DeepSeek V3 still outperforms LLaMA-4, reaching a tile value of 256, whereas LLaMA-4 only reaches 128, the same as a random move algorithm. In Candy Crush and Sokoban, which demand spatial reasoning capability, LLaMA 4 struggles and fails to crack the first level in both 🧩. Checkout our leaderboard for more detailed results!

Hao AI Lab's profile picture
Hao AI Lab1 year ago

GameArena is hard to overfit, our gaming benchmark features: 1. A transparent model evaluation paradigm using classic games, ensuring reproducibility by anyone. 2. Strong robustness to data contamination, thanks to the large number of games and the intricate nature of gaming environments.

Hao AI Lab's profile picture
Hao AI Lab1 year ago

Check out our repo and explore how video games can push the boundaries of AI capability. A new video game is also on the way 👾🕹️, and we’re looking forward to watching LLaMA-4 behemoth shakes up the leaderboard. 🔗

Ithy's profile picture
Ithy1 year ago

What happens when you combine every AI? It's time for something better than ChatGPT...

Related Videos