
Arena.ai
@arena • 219,434 subscribers
Where AI meets the real world. We measure and advance the frontier of AI through community-driven evaluation. We’re hiring → https://t.co/XBZCrsdD77
Shorts
Videos

Introducing Agent Mode: Agentic AI is now measured in the Arena. Agent Mode can do deep research, create reports, generate images, build websites, debug code, and more. It completes more complex tasks by using tools like web search, bash in a sandbox environment, image generation, file writing, and asking follow-up questions. Frontier models are waiting for you in Agent Mode to take on real-world tasks. GPT-5.5, Claude Opus 4.7, Gemini 3.1 Pro, and top open models. Test them yourself.
Arena.ai450,204 Aufrufe • vor 3 Monaten

Code Arena just leveled up with fullstack capabilities 🚀 Introducing the new Fullstack Code Arena. We’re moving beyond frontend prototypes to fullstack development complete with databases, API keys, and fast deployments. Build, iterate, and ship real-world software — all in one place. Models now act as agents in the Code Arena, using structured tool calls to plan, execute, and refine in real time with real world tasks. Read more about it in the thread 🧵
Arena.ai285,064 Aufrufe • vor 2 Monaten

During an agentic session, a quick “hi” after a coffee break can cost a full dollar. But it’s not because of the “hi.” It’s the million tokens of context you’re repaying for the moment your cache goes cold. Here's the caching math from Evan: 0:00 Why you're billed for last turn's context, not just this one 0:17 Cache hits: ~10% of full price 0:41 Agentic AI = way more back-and-forth than chat 1:16 Paying full price for context even on a 1-token tool call 1:40 Cache hit vs. cache miss 2:13 The $1 "hi": walk away for an hour, come back to a cold cache 2:59 Cost grows by the square, even with caching 3:28 How context compaction helps 4:03 Why Claude Code/Codex may compact at ~200-300K, not the full 1M window
Arena.ai51,694 Aufrufe • vor 15 Tagen

Arena reached a $100M annual revenue run rate just 8 months after launching our evaluation product. We started as a research project at UC Berkeley with a simple mission: measure AI progress through real-world use. As AI shifts from chatbots to agents taking on longer, higher-stakes work, the problem matters more than ever. Today, Arena measures real-world AI utility with a community of tens of millions. With Agent Arena, we’re evaluating long-running agents on complex, real-world tasks - how they use tools, adapt to feedback, recover from errors, and accomplish goals set by humans. We are excited to keep deepening our work in agentic evaluations. Here’s Anastasios Nikolas Angelopoulos on what this milestone means and where we go from here:
Arena.ai211,479 Aufrufe • vor 2 Monaten

Agent Categories and Costs are now in the Agent Leaderboard! Two questions guide model selection: task-specific performance and cost to completion. Based on analysis of 1.7M+ real-world Agent Mode sessions, the leaderboard now lets users: - Filter the leaderboard by Code, Work, and Chat to compare model performance by task type - View real cost-per-task for each model, including the Pareto frontier of the most cost-efficient options at each performance level
Arena.ai29,227 Aufrufe • vor 15 Tagen

Arena Conversations: Training Laguna with Poolside researchers, Connor Adams and Aalhad Patankar. With Peter Gostev (SF 24-28 August), they dig into pre-training that builds broad capability, and post-training to reveal whether those skills actually work together across 250,000 long-horizon trajectories. Check out the full conversation on our YouTube to hear how Poolside is pairing humans, agents, and evidence-annotated reviews to sharpen future training loops. Link below.
Arena.ai16,171 Aufrufe • vor 11 Tagen

Arena intern and UCLA PhD candidate, Hengguang Zhou, introduces Trace-and-Amplify (TA), a framework for collecting training-time reward-hacking trajectories at scale without hacking instructions. Monitors trained and evaluated on prompt-elicited hacking trajectories can achieve high detection accuracy, but often fail to transfer to training-time reward-hacking trajectories that emerge during RL without hacking instructions. Trace-and-Amplify enables scalable collection of these training-time trajectories, producing monitors that generalize much better to real and held-out hacking types. Detection accuracy 59.98% (PE-trained) → 90.16% (TA-trained) compared to 97.1% on prompted hacks → 28.0% on training-time hacks. 0:00 – OpenAI's ExploitGym cyberattack benchmark exploit 1:04 – Goodhart's Law and the CoastRunners boat-racing hack (2016) 2:04 – Gaming the evaluator: the robot-hand grasping example (2017) 3:10 – Reward hacking in code generation: hard-coding, test-rewriting, skipping eval 4:20 – A standard defense: reward-hacking monitors 4:58 – Monitor architectures: zero-shot LLMs, fine-tuned BERT, hidden-state probes 6:11 – Where monitor training data comes from today: prompted hacks 7:03 – The core question: do prompted hacks represent real hacks? 7:35 – Why this matters: RL post-training is the standard recipe for frontier models 8:20 – Why nobody's checked this before (hacking is rare, labeling isn't scalable) 9:23 – Introducing the method: Trace-and-Amplify 9:49 – The Tracer: a contradictory unit test that locates evaluation-gaming 10:50 – Amplify: collecting hacking rollouts at scale during RL training 11:32 – Experiment setup: Qwen2.5-Coder, DeepSeek-Coder, LeetCode/TACO 12:16 – Finding #1: prompt-trained monitors don't transfer to real hacks 14:40 – Can strong zero-shot judges (GPT-4.1, o4-mini) do better? 15:48 – Finding #2: monitors trained on real hacks generalize much better to unseen hacks 17:04 – Ruling out artifacts introduced by the method 17:55 – Why the gap? Real hacking is more hidden than prompted hacking 20:05 – Three takeaways, limitations, and future work
Arena.ai33,419 Aufrufe • vor 25 Tagen

Kimi K3 passed Fable 5 on Code Arena just 6 weeks after Fable's release, landing #1 in 6 of 7 frontend domains. We sat down with Arena CEO and co-founder Anastasios Nikolas Angelopoulos for a rapid fire conversation on what the data from real-world use tells us 👇 0:00 Biggest release of the year, or overreaction? 0:33 Kimi K3 beats Fable on Code Arena: verifying the votes 1:35 1679 points: a 17-place jump from K2.6 2:37 #1 in 6 of 7 frontend domains — but still #2 in gaming 3:15 Permanent shift or snapshot? Labs distilling Chinese models 4:23 Regulating Chinese models: the case for and against 6:30 $3 in, $15 out: open source is no longer 10% of the price 7:30 If you're Anthropic or OpenAI, what's your move? 8:09 How long before the leaderboard flips again?
Arena.ai50,961 Aufrufe • vor 1 Monat

📢We’re excited to share that we’ve raised $100M in seed funding to support LMArena and continue our research on reliable AI. Led by a16z and UC Investments (University of California), we're proud to have the support of those that believe in both the science and the mission. We’re focused on building a neutral, open, community-driven platform that helps the world understand and improve the performance of AI models on real queries from real users. Also, big news is coming next week!👀 We're relaunching LMArena with a whole new look built directly with community feedback from the ground up 🧱 Link in thread.
Arena.ai436,189 Aufrufe • vor 1 Jahr

GPT-5.2-high by OpenAI is off to a strong start in the Code Arena. ⚡️ If you’re new here: Code Arena is where AI models build full web apps, tools, and interactive sites — all from a single prompt. Watch the video to see GPT-5.2-high in action, then try your own prompt and reply with your creation below. ⬇️
Arena.ai246,430 Aufrufe • vor 8 Monaten

Arena Conversations: Turning tens of thousands of experiments across data mixes, into models that move the frontier. Peter Gostev (SF 24-28 August) talks with two researchers Connor Adams and Aalhad Patankar from Poolside about how they learn quickly through iteration, massive pre-training runs, RL decisions, and more. Watch the full conversation on our YouTube. Find the link below.
Arena.ai14,068 Aufrufe • vor 12 Tagen

Arena Conversations: Where’s the line between resourceful agent behavior and reward hacking? Poolside researchers, Connor Adams and Aalhad Patankar, discuss benchmark awareness, instruction following, and where persistence starts to look like misalignment. And to hear about Peter Gostev (SF 24-28 August)'s agent sending an email on his behalf without request… Check out the full interview below.
Arena.ai14,793 Aufrufe • vor 13 Tagen

Arena AutoEval can quickly evaluate new models, allowing model labs to leverage results faster. It compresses the evaluation feedback loop from weeks to hours, enabling rapid, human-preference-based insight into a model’s capabilities. For more details, see the full video in the post below.
Arena.ai21,456 Aufrufe • vor 21 Tagen

Our CEO and Co-Founder, Anastasios Nikolas Angelopoulos, joined Harry Stebbings on 20VC to share his thoughts on some of the most consequential technical, economic, and security questions shaping AI: - U.S. vs China competition in frontier and open-weight models - The future of the U.S. open-source AI ecosystem - Kimi K3 release and what it signals about China’s model capabilities - Gravity of the recent cyber security incidents - Which neo-labs will breakthrough into durable businesses - The next generation of business moats: proprietary data, network effects, and ownership of the intelligence stack Watch the full podcast below.
Arena.ai26,278 Aufrufe • vor 1 Monat

More from Opus 5 and Peter Gostev. Put Opus 5 to the test with your own prompts in Agent Mode!
Arena.ai29,222 Aufrufe • vor 1 Monat

The NEW LMArena is officially live! 🎉 ✨ New Logo! ⚡️ Better, faster UI/UX for chat and leaderboard 📱 Mobile optimized 💬 Chat history 🧭 Clearer leaderboard navigation 🤖 Many modalities in one place: vision, image, and more coming soon Try it now at lmarena dot ai! (Link in 🧵)
Arena.ai267,006 Aufrufe • vor 1 Jahr

We put the top three Code Arena models head-to-head: Opus 4.5 Thinking 32k, Opus 4.5, and Gemini 3 Pro. They’re just 20 points apart. Same tough prompts, different results. Here’s what stood out. Remember, your votes drive the rankings. Watch how these contenders move on the leaderboard as more votes come in. Check out the comparisons in the thread below. 🧵
Arena.ai163,929 Aufrufe • vor 9 Monaten