正在加载视频...
视频加载失败
Introducing Agent Mode: Agentic AI is now measured in the Arena. Agent Mode can do deep research, create reports, generate images, build websites, debug code, and more. It completes more complex tasks by using tools like web search, bash in a sandbox environment, image generation, file writing, and asking... show more
504,119 次观看 • 3 个月前 •via X (Twitter)
37 条评论

Agent Mode helps streamline multi-step workflows, minimizing the need for multiple prompts. Learn how you can evaluate the frontier of AI:

Read more about Agent Mode, dig into the FAQ, and get a preview of what we've learned so far on our blog at:

Start evaluating agentic AI on Arena today with Agent Mode at:

Why gemini-3-pro-image-preview-2k (nano-banana-pro) can't be used now, please bring back gemini-3-pro-image-preview-2k (nano-banana-pro).

Finally being able to benchmark agentic AI on real world tasks instead of just chat is what the evals space actually needed

Sounds like a big step toward full AI agents.

finally evaluating actual agentic workflows.

agent mode is where the puck is going. the magic isn't a smarter chat box — it's the agent taking 20 steps on its own and handing you a finished result. been shipping this exact loop in the browser, no setup. it changes how people expect software to feel.

finally my time to shine

I hope Arena AI will be available as app for Android.

@arena cool stuff! curious how it handles debugging. does it spot errors or suggest fixes too?

Agent Mode woke up and chose everyone's job description. 🤯🔥

Agent Mode gets more useful when the trajectory is inspectable, not just the final score. For builders, the eval should answer: did the agent recover from bad tool outputs, choose sane handoff points, and leave an audit trail a human can trust?

This kind of evaluation is needed because agents are hard to judge from demos. For everyday workflows, I care about 3 things: can it finish the task, can I inspect the steps, and can I recover when it goes wrong. That is the difference between impressive and actually useful.

.

@Apple @jadi

Useful leaderboard work. One missing layer for Agent Mode: process stability across change. Not only whether a model wins isolated comparisons, but whether it preserves goals, constraints, corrections, prior decisions and state across evolving tasks without drift. That should be measured publicly.

Agent Mode is a big step toward practical AI workflows. Excited to see frontier models handling real-world tasks in one unified experience.

Agent Mode brings AI closer to real-world problem solving with powerful tools and capabilities.

Curious how the sandbox metrics translate to production workloads, seeing real world latency will be key

agent arena is what the space needed. everyone claims their agent is the best, now there's a way to actually measure it

Chế độ Agent: AI Agentic là một công cụ mạnh mẽ cho nghiên cứu và xử lý nhiệm vụ phức tạp. Hãy thử ngay!

Fix your broken captcha

Agentic AI in the Arena is a huge step forward. Exciting times.

A significant step toward evaluating AI on real world capabilities

Feature request: please add “Pin” 📌 and “Delete” 🗑️ for Arenas. Pin important Arenas to the top, and Delete permanently removes unused ones. Rename + Archive are great—these would make Arena management much better.

👍

Fix your broken captcha!

Measuring agentic AI is the hard problem nobody talks about enough. Single-turn benchmarks are dead — what matters now is whether a model can navigate multi-step tasks without hallucinating itself into a corner. Agent Mode leaderboards are the future of model eval.

Change name from Agent to one of: GLITCH: Gremlins Lurking In The Code Here SNAP: Something Not Acting Perfectly FLOP: Failed, Load Over Please OOPS: Operation Obstructed, Please Submit (Again)

你们最好是能给我快一点测试出 glm5.3 的绑定

What on world is battle mode 😭

Probably the best app in recent times

لدي مشكلة قمت بعمل خلال يومين وتم إنتاج ٦٥ مستند في محادثة للأسف قبل أن اسجل دخول وعندما قمت بتسجيل الدخول اختفت المحادثة الرجاء مساعدتي لاستعادة المحادثة والملفات المفقودة

Hi, is there an issue with Agent Mode right now? It seems to be stuck and not responding to prompts. Is there any ongoing maintenance or outage?

Agent Mode is the right direction. Testing models on real multi-step tasks instead of just chat is much more useful.

🚀 Agentic AI is moving from simple chat to real execution. Excited to see where this goes!

