Video yükleniyor...
Video Yüklenemedi
Nex-N2.5 feels less like “another model launch” and more like a push toward agents that can actually finish real work. The part that stands out to me is Nex-N2.5 Pro: a 397B multimodal model built for Vision, Computer Use, and long-horizon interaction. Instead of only understanding a screen, it... show more
66,447 görüntüleme • 6 gün önce •via X (Twitter)
33 Yorum

This feels less like another model drop and more like a push toward agents that finish the job. Reading a changing UI and staying in the software is the part that stands out.

The shift from AI that gives instructions to AI that can actually operate the tools is massive. This is the kind of Computer Use demo worth watching

Real diff is reading UI changes, not just chat.

That shift from writing a code snippet and stopping to maintaining a live visual feedback loop is what finally makes "computer use" feel practical rather than like a heavily scripted demo.

It can navigate changing interfaces and correct itself, which specific software workflow are you planning to hand over to it first

Free tier rate-limited hard? Wanna run overnight.

Idea to a Blender scene, then check the screen, fix it, verify again. That loop is the useful claim. Not a one-shot script.

Having models like Pro handle UI-heavy tools, CAD, and desktop software natively means agents can finally bridge the gap between abstract user intent and a verified, working outcome.

Games, Blender, CAD, browsers. The brief is pointing at desktop work, not chat demos. Long-horizon interaction is the bar.

Mini and Pro for multimodal screen work, Max for the heavy reasoning jobs. Same-day access is how you actually test that split.

A 397B model that can keep track of changing screens while working through a long task is definitely something I want to stress-test.

The shift from “describe the click” to “keep going when the UI changes” is the only claim that matters. Curious which of those demos was a live loop vs a cleaned-up recording — Blender and FreeCAD are very different failure modes.

Free on OpenRouter for a limited time makes this even easier to experiment with. I’d start with Pro and throw a real Blender or browser workflow at it.

Most demos die after 1 action. This one navigates, battles, recovers. Actually useful. Free on OpenRouter 🔥

This is the real test for AI agents. Playing Pokémon for hundreds of steps is insane. Trying it on OpenRouter now

The shift from answering to actually operating software is the part that matters. 👀 That’s where AI starts feeling like an agent instead of a chatbot.

写完代码还会自己跑一遍、看界面找问题再修,这段确实很戳人

Splitting the family into multimodal vision variants for GUI interaction (Mini and Pro) alongside a massive text-only MoE powerhouse (Max) gives developers the right scaling axis depending on whether the bottleneck is visual tracking or deep reasoning.

This is the real leap, agents that actually open the tools and finish the work.

视频里最有说服力的不是会点按钮,而是它能跟着界面变化继续往下做:从零件摆放到灯光、分层、试图,中间还在改脚本。现在多数 Computer Use 还停在‘看懂屏幕’,真正卡在长程反馈闭环。Nex-N2.5 Pro 如果能在 Blender / FreeCAD 这类工具上稳定跑完一轮,比再堆一个基准分数更有意义。准备去 OpenRouter 试 Mini/Pro,想先看它处理报错后的自我修正稳不稳。

Agent that fixes its own bug > bigger params imo.

Free trial first; paid only if it survives long runs.

真正拉开差距的不是参数,是能盯着界面变化把活干完。

reading a changing UI and staying in the software is the part that stands out to me

397B built for vision and computer use only counts if it can stay in the tool step by step. Most models stop after the first pretty output.

this feels less like another model drop and more like agents that can actually finish the job

After writing the code, they even run it themselves once, check the interface for issues, and fix them—this part really hits home.

There lies the real test for AI agents. Playing Pokémon for many numbers of steps is insane. Trying it on OpenRouter now

Just tried Nex-N2.5 Pro on my Blender gear assembly. It set cameras, fixed lighting, organized parts, and rendered a clear exploded view on its own. First model that actually finishes the desktop work instead of only describing steps.

Mini and Pro for multimodal computer use, Max at 1.6T for long reasoning. The split is clearer than another single flagship drop.

Watching UI states makes long tasks practical

This is huge "Do it" agents "talk about it" agents

The CAD and Blender examples are the part I’d watch. Once an agent can stay inside professional software long enough to finish a task, the use case gets a lot more serious.

