Arena.ai's banner
Arena.ai's profile picture

Arena.ai

@arena219,434 subscribers

Where AI meets the real world. We measure and advance the frontier of AI through community-driven evaluation. We’re hiring → https://t.co/XBZCrsdD77

Shorts

Arena Trends: Text-to-Image, Jan 2026 – Apr 2026 For most of the year, Google DeepMind and OpenAI traded the top spot within a tight margin - GPT-Image vs. Nano Banana - with the rest of the field clustered below 1,200. Today, GPT-Image-2 breaks away with a score of 1,512, 242 points ahead of #2 Google. The frontier continues to move.

Arena Trends: Text-to-Image, Jan 2026 – Apr 2026 For most of the year, Google DeepMind and OpenAI traded the top spot within a tight margin - GPT-Image vs. Nano Banana - with the rest of the field clustered below 1,200. Today, GPT-Image-2 breaks away with a score of 1,512, 242 points ahead of #2 Google. The frontier continues to move.

515,204 Aufrufe

A new episode of Arena Conversations is dropping tomorrow at 9 AM ET / 6 AM PT: Featuring Bryan Catanzaro, the VP of Applied Deep Learning Research at NVIDIA AI! With Peter Gostev (SF 24-28 August), they unpack the role of specialized teacher models in building stronger capabilities, the growing inference challenge in post-training, the importance of open source models, and what it all means for NVIDIA and how they build Nemotron. Stay tuned tomorrow, August 24th.

A new episode of Arena Conversations is dropping tomorrow at 9 AM ET / 6 AM PT: Featuring Bryan Catanzaro, the VP of Applied Deep Learning Research at NVIDIA AI! With Peter Gostev (SF 24-28 August), they unpack the role of specialized teacher models in building stronger capabilities, the growing inference challenge in post-training, the importance of open source models, and what it all means for NVIDIA and how they build Nemotron. Stay tuned tomorrow, August 24th.

33,289 Aufrufe

US vs China update. Stanford's AI Index put the US–China gap at 2.7%. Here's what two years of real-world use from the Text Arena shows. Gap three years ago: +278. Today: +29. Anthropic's Claude Opus 4.6 Thinking vs. Baidu's ERNIE for Developers Ernie 5.1 at the top. The US has never lost #1, but the race keeps closing.

US vs China update. Stanford's AI Index put the US–China gap at 2.7%. Here's what two years of real-world use from the Text Arena shows. Gap three years ago: +278. Today: +29. Anthropic's Claude Opus 4.6 Thinking vs. Baidu's ERNIE for Developers Ernie 5.1 at the top. The US has never lost #1, but the race keeps closing.

58,928 Aufrufe

We decided to take Paul Jankura’s Claude Opus 4.5 out for a test drive vs. the current #1 ranking model in Code Arena: Gemini 3 Pro. Same prompt, different outputs. Let’s take a look. Remember, your votes drive the leaderboards. We’ll see how Claude Opus 4.5 stacks up in the coming days! Check out some of the comparisons, like how Claude Opus 4.5 handled the “Pyramids of Giza” prompt, in thread. 🧵

We decided to take Paul Jankura’s Claude Opus 4.5 out for a test drive vs. the current #1 ranking model in Code Arena: Gemini 3 Pro. Same prompt, different outputs. Let’s take a look. Remember, your votes drive the leaderboards. We’ll see how Claude Opus 4.5 stacks up in the coming days! Check out some of the comparisons, like how Claude Opus 4.5 handled the “Pyramids of Giza” prompt, in thread. 🧵

88,977 Aufrufe

📊Arena Trend update for August 2024 - Feb 2025: After a few DeepSeek jumps last month, xAI leaps forward to the top of the leaderboard. The AI race continues! 📈 animation credit: Peter Gostev

📊Arena Trend update for August 2024 - Feb 2025: After a few DeepSeek jumps last month, xAI leaps forward to the top of the leaderboard. The AI race continues! 📈 animation credit: Peter Gostev

154,425 Aufrufe

Arena leaderboards now include Price and Context. - Price is shown as input / output cost per 1M tokens, and context shows the maximum context window. Compare Arena scores based on what matters for your use case.

Arena leaderboards now include Price and Context. - Price is shown as input / output cost per 1M tokens, and context shows the maximum context window. Compare Arena scores based on what matters for your use case.

28,517 Aufrufe

Created by Gemini 3 Pro in one shot!

Created by Gemini 3 Pro in one shot!

36,569 Aufrufe

📈Arena Trends Update We pulled Arena scores for the Top 10 labs in Text for the past 6 months (Sept-2025-Feb 2026), and the competitive spread is shifting again. With tighter confidence intervals and new entries in the mix, the frontier continues to shift. Stay tuned for more insights as we dive deeper into the top open models for February later this week. Let us know what you found the most surprising in the comments. 👇

📈Arena Trends Update We pulled Arena scores for the Top 10 labs in Text for the past 6 months (Sept-2025-Feb 2026), and the competitive spread is shifting again. With tighter confidence intervals and new entries in the mix, the frontier continues to shift. Stay tuned for more insights as we dive deeper into the top open models for February later this week. Let us know what you found the most surprising in the comments. 👇

23,239 Aufrufe

🍌 Thousands of new people jumped into Image Arena Battle mode this week - our intern can barely keep up! What happens in Battle mode? 🧵 We partner directly with model providers to give you early access to cutting-edge models still in development, often before you can try them anywhere else. These pre-release models are tested in Battle mode. Details in the thread 👇

🍌 Thousands of new people jumped into Image Arena Battle mode this week - our intern can barely keep up! What happens in Battle mode? 🧵 We partner directly with model providers to give you early access to cutting-edge models still in development, often before you can try them anywhere else. These pre-release models are tested in Battle mode. Details in the thread 👇

16,938 Aufrufe

Videos

arena's profile picture

Arena intern and UCLA PhD candidate, Hengguang Zhou, introduces Trace-and-Amplify (TA), a framework for collecting training-time reward-hacking trajectories at scale without hacking instructions. Monitors trained and evaluated on prompt-elicited hacking trajectories can achieve high detection accuracy, but often fail to transfer to training-time reward-hacking trajectories that emerge during RL without hacking instructions. Trace-and-Amplify enables scalable collection of these training-time trajectories, producing monitors that generalize much better to real and held-out hacking types. Detection accuracy 59.98% (PE-trained) → 90.16% (TA-trained) compared to 97.1% on prompted hacks → 28.0% on training-time hacks. 0:00 – OpenAI's ExploitGym cyberattack benchmark exploit 1:04 – Goodhart's Law and the CoastRunners boat-racing hack (2016) 2:04 – Gaming the evaluator: the robot-hand grasping example (2017) 3:10 – Reward hacking in code generation: hard-coding, test-rewriting, skipping eval 4:20 – A standard defense: reward-hacking monitors 4:58 – Monitor architectures: zero-shot LLMs, fine-tuned BERT, hidden-state probes 6:11 – Where monitor training data comes from today: prompted hacks 7:03 – The core question: do prompted hacks represent real hacks? 7:35 – Why this matters: RL post-training is the standard recipe for frontier models 8:20 – Why nobody's checked this before (hacking is rare, labeling isn't scalable) 9:23 – Introducing the method: Trace-and-Amplify 9:49 – The Tracer: a contradictory unit test that locates evaluation-gaming 10:50 – Amplify: collecting hacking rollouts at scale during RL training 11:32 – Experiment setup: Qwen2.5-Coder, DeepSeek-Coder, LeetCode/TACO 12:16 – Finding #1: prompt-trained monitors don't transfer to real hacks 14:40 – Can strong zero-shot judges (GPT-4.1, o4-mini) do better? 15:48 – Finding #2: monitors trained on real hacks generalize much better to unseen hacks 17:04 – Ruling out artifacts introduced by the method 17:55 – Why the gap? Real hacking is more hidden than prompted hacking 20:05 – Three takeaways, limitations, and future work

Arena.ai

33,419 Aufrufe • vor 25 Tagen