Loading video...
Video Failed to Load
LLaMA-4 Maverick performs well on reasoning benchmarks and ranks 2nd on the Chatbot Arena, yet its true performance remains controversial. What if we put them in a transparent gaming environment? 🎮 Our benchmark tells a different story...🤔 Will true intelligence shine through play? Let’s find out 👇
39,141 views • 1 year ago •via X (Twitter)
4 Comments

We directly compared LLaMA-4 Maverick with its reported peers. Despite of LLaMA-4’s strong performance on static benchmarks, in 2048, DeepSeek V3 still outperforms LLaMA-4, reaching a tile value of 256, whereas LLaMA-4 only reaches 128, the same as a random move algorithm. In Candy Crush and Sokoban, which demand spatial reasoning capability, LLaMA 4 struggles and fails to crack the first level in both 🧩. Checkout our leaderboard for more detailed results!

GameArena is hard to overfit, our gaming benchmark features: 1. A transparent model evaluation paradigm using classic games, ensuring reproducibility by anyone. 2. Strong robustness to data contamination, thanks to the large number of games and the intricate nature of gaming environments.

Check out our repo and explore how video games can push the boundaries of AI capability. A new video game is also on the way 👾🕹️, and we’re looking forward to watching LLaMA-4 behemoth shakes up the leaderboard. 🔗

What happens when you combine every AI? It's time for something better than ChatGPT...
