正在加载视频...

视频加载失败

Introducing Agent Mode: Agentic AI is now measured in the Arena. Agent Mode can do deep research, create reports, generate images, build websites, debug code, and more. It completes more complex tasks by using tools like web search, bash in a sandbox environment, image generation, file writing, and asking...

504,119 次观看 • 3 个月前 •via X (Twitter)

37 条评论

Arena.ai 的头像
Arena.ai3 个月前

Agent Mode helps streamline multi-step workflows, minimizing the need for multiple prompts. Learn how you can evaluate the frontier of AI:

Arena.ai 的头像
Arena.ai3 个月前

Read more about Agent Mode, dig into the FAQ, and get a preview of what we've learned so far on our blog at:

Arena.ai 的头像
Arena.ai3 个月前

Start evaluating agentic AI on Arena today with Agent Mode at:

Didaa 的头像
Didaa3 个月前

Why gemini-3-pro-image-preview-2k (nano-banana-pro) can't be used now, please bring back gemini-3-pro-image-preview-2k (nano-banana-pro).

AI Mastery Guide 的头像
AI Mastery Guide3 个月前

Finally being able to benchmark agentic AI on real world tasks instead of just chat is what the evals space actually needed

Aaliya 的头像
Aaliya3 个月前

Sounds like a big step toward full AI agents.

UmarAi 的头像
UmarAi3 个月前

finally evaluating actual agentic workflows.

BOB CHEN 的头像
BOB CHEN3 个月前

agent mode is where the puck is going. the magic isn't a smarter chat box — it's the agent taking 20 steps on its own and handing you a finished result. been shipping this exact loop in the browser, no setup. it changes how people expect software to feel.

Wes Sander 的头像
Wes Sander3 个月前

finally my time to shine

Sema Rose 的头像
Sema Rose3 个月前

I hope Arena AI will be available as app for Android.

Hussain Hashim | Building SundayBack 的头像
Hussain Hashim | Building SundayBack3 个月前

@arena cool stuff! curious how it handles debugging. does it spot errors or suggest fixes too?

XYBER X 的头像
XYBER X1 个月前

Agent Mode woke up and chose everyone's job description. 🤯🔥

Mart 的头像
Mart3 个月前

Agent Mode gets more useful when the trajectory is inspectable, not just the final score. For builders, the eval should answer: did the agent recover from bad tool outputs, choose sane handoff points, and leave an audit trail a human can trust?

Solomon Omolabi 的头像
Solomon Omolabi3 个月前

This kind of evaluation is needed because agents are hard to judge from demos. For everyday workflows, I care about 3 things: can it finish the task, can I inspect the steps, and can I recover when it goes wrong. That is the difference between impressive and actually useful.

k.aqh 的头像
k.aqh2 个月前

.

MOHAMMAD 的头像
MOHAMMAD3 个月前

@Apple @jadi

Eduardo Rivera | Continuidad 的头像
Eduardo Rivera | Continuidad2 个月前

Useful leaderboard work. One missing layer for Agent Mode: process stability across change. Not only whether a model wins isolated comparisons, but whether it preserves goals, constraints, corrections, prior decisions and state across evolving tasks without drift. That should be measured publicly.

Blake Thompson 的头像
Blake Thompson2 个月前

Agent Mode is a big step toward practical AI workflows. Excited to see frontier models handling real-world tasks in one unified experience.

Blake Thompson 的头像
Blake Thompson2 个月前

Agent Mode brings AI closer to real-world problem solving with powerful tools and capabilities.

Shubham Sharma | AI & Tech 的头像
Shubham Sharma | AI & Tech3 个月前

Curious how the sandbox metrics translate to production workloads, seeing real world latency will be key

Volodymyr Pavlenko 的头像
Volodymyr Pavlenko3 个月前

agent arena is what the space needed. everyone claims their agent is the best, now there's a way to actually measure it

Terry Bui 的头像
Terry Bui3 个月前

Chế độ Agent: AI Agentic là một công cụ mạnh mẽ cho nghiên cứu và xử lý nhiệm vụ phức tạp. Hãy thử ngay!

Juusepson64 的头像
Juusepson642 个月前

Fix your broken captcha

RAZA | AI EXPLORER 的头像
RAZA | AI EXPLORER15 天前

Agentic AI in the Arena is a huge step forward. Exciting times.

ZenithAi 的头像
ZenithAi1 个月前

A significant step toward evaluating AI on real world capabilities

Washim Reja 的头像
Washim Reja8 天前

Feature request: please add “Pin” 📌 and “Delete” 🗑️ for Arenas. Pin important Arenas to the top, and Delete permanently removes unused ones. Rename + Archive are great—these would make Arena management much better.

Muhammad 的头像
Muhammad3 个月前

👍

Juusepson64 的头像
Juusepson642 个月前

Fix your broken captcha!

LoongLab 🐉 的头像
LoongLab 🐉3 个月前

Measuring agentic AI is the hard problem nobody talks about enough. Single-turn benchmarks are dead — what matters now is whether a model can navigate multi-step tasks without hallucinating itself into a corner. Agent Mode leaderboards are the future of model eval.

John R. 的头像
John R.2 个月前

Change name from Agent to one of: GLITCH: Gremlins Lurking In The Code Here SNAP: Something Not Acting Perfectly FLOP: Failed, Load Over Please OOPS: Operation Obstructed, Please Submit (Again)

王培宇 的头像
王培宇1 个月前

你们最好是能给我快一点测试出 glm5.3 的绑定

virat mankali 的头像
virat mankali3 个月前

What on world is battle mode 😭

Boris AI 的头像
Boris AI2 个月前

Probably the best app in recent times

Mahmoud 的头像
Mahmoud3 个月前

لدي مشكلة قمت بعمل خلال يومين وتم إنتاج ٦٥ مستند في محادثة للأسف قبل أن اسجل دخول وعندما قمت بتسجيل الدخول اختفت المحادثة الرجاء مساعدتي لاستعادة المحادثة والملفات المفقودة

SmartSignal Pro 的头像
SmartSignal Pro15 天前

Hi, is there an issue with Agent Mode right now? It seems to be stuck and not responding to prompts. Is there any ongoing maintenance or outage?

Hex Nova 的头像
Hex Nova1 个月前

Agent Mode is the right direction. Testing models on real multi-step tasks instead of just chat is much more useful.

Hamid AI 的头像
Hamid AI2 个月前

🚀 Agentic AI is moving from simple chat to real execution. Excited to see where this goes!

相关视频