Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

AI agents are moving past the point where generating an answer is enough. The real question is whether an agent can stay inside a real environment long enough to actually finish the job. That’s what makes Nex-N2.5 interesting to me It is built around long-horizon computer use, with agents...

20,759 görüntüleme • 8 gün önce •via X (Twitter)

35 Yorum

Aureus profil fotoğrafı
Aureus8 gün önce

This is the part I actually care about. Not another first-try answer, an agent that stays in the tool, sees the result, and keeps fixing until the work is finished.

Synapse profil fotoğrafı
Synapse8 gün önce

The SQL interpreter example is the right kind of test. First output is easy. Staying in the environment long enough to add WHERE, JOIN, tests, and fixes is the actual work.

maximillion retardio profil fotoğrafı
maximillion retardio8 gün önce

The write → run → inspect → fix loop is the part that matters. Generating a SQL interpreter once is easy. Staying in the environment long enough to add JOIN, aliases, and tests is the actual test.

Almusty profil fotoğrafı
Almusty8 gün önce

Building a working SQL interpreter step-by-step through self-testing is a massive flex for a long-horizon agent. Way more impressive than just spitting out a static script on the first try.

Almusty profil fotoğrafı
Almusty8 gün önce

The write-run-fix cycle is where developers spend most of their time anyway, so having an agent that can handle that loop autonomously is a massive time-saver.

Ryvox profil fotoğrafı
Ryvox8 gün önce

generating an answer isn’t the bar anymore. staying in the environment long enough to finish the job is.

Equinox profil fotoğrafı
Equinox8 gün önce

Write → run → inspect → fix is a cleaner way to judge these models than a benchmark screenshot. If the agent drops out after the first artifact, it is still just a generator.

Almusty profil fotoğrafı
Almusty8 gün önce

That closed-loop execution is honestly the real breakthrough. Most models just hand you code that breaks instantly, so watching an agent actually test and debug its own SQL logic changes everything.

Almusty profil fotoğrafı
Almusty8 gün önce

Perfectly put. We're finally shifting away from glorified autocomplete into actual autonomous execution that can see a multi-step task through to the end.

Starlit_void profil fotoğrafı
Starlit_void8 gün önce

Visual feedback plus self correction is the piece people keep asking for. An agent that can see what broke and keep iterating is more useful than one that only writes code once.

Em Kei profil fotoğrafı
Em Kei8 gün önce

This is the part of agentic AI I find most interesting: not just generating an answer, but being able to interact with the environment, check the result, and keep improving until the task is actually done.

Em Kei profil fotoğrafı
Em Kei8 gün önce

The real test for AI agents isn’t how impressive the first response looks. It’s what happens when things go wrong. Being able to detect an error, adapt, and keep going is where agentic systems can become genuinely powerful.

Zyron profil fotoğrafı
Zyron8 gün önce

The feedback loop part is the real upgrade. Write → run → fix → verify is what I actually want from agents.

daemon profil fotoğrafı
daemon8 gün önce

Visual feedback matters more than people admit. An agent that cannot see the result is guessing. An agent that can inspect the run and patch the next piece is doing work.

Em Kei profil fotoğrafı
Em Kei8 gün önce

The write → run → inspect → fix → verify loop feels like a much more meaningful benchmark for AI agents than simply asking whether the first output looks good. That’s where things start getting really interesting.

MavisKrypt ⚖️ 🦔 profil fotoğrafı
MavisKrypt ⚖️ 🦔8 gün önce

The real upgrade is the feedback loop, not just stronger outputs. Agents that can test, inspect, correct, and verify are much closer to actually completing tasks.

useetek profil fotoğrafı
useetek8 gün önce

Many of the most important contributions happen quietly before the community sees the results.

Nova profil fotoğrafı
Nova8 gün önce

The clip is doing the talking. You can watch it work inside the software, check the view, correct the path, and keep going instead of stopping at the first pass.

Zentry profil fotoğrafı
Zentry8 gün önce

this is the part that actually matters. generating the first draft is easy. staying in the loop and finishing the job is the hard part.

0xUpkiba profil fotoğrafı
0xUpkiba8 gün önce

Sounds like Nex-N2.5 has some serious potential, bro! Excited to see where it goes.

sulaiman Waleed profil fotoğrafı
sulaiman Waleed8 gün önce

Build SQL → test → find gaps → implement. WHERE, JOIN, subqueries handled. This is long-horizon Computer Use 👀

smalldas profil fotoğrafı
smalldas8 gün önce

This is the future of AI—no more good enough, just relentless iteration until it’s actually done. NexN2.5 doing SQL dev like a real engineer is wild. Hope they don’t break my laptop in the process.

Joe profil fotoğrafı
Joe8 gün önce

I’m more interested in crypto when the technology is solving real problems instead of chasing hype.

Ariel Lillie profil fotoğrafı
Ariel Lillie8 gün önce

long horizon browser use and visual feedback sounds key for finishing

HDUX profil fotoğrafı
HDUX8 gün önce

The write-run-fix loop is exactly where I've seen most agents fall apart. They dump code and bail. Curious how Nex-N2.5 handles edge cases when the SQL interpreter hits something it's never seen before though. Self-testing is cool until it gets stuck in a loop.

Crypto Sapiens profil fotoğrafı
Crypto Sapiens8 gün önce

Longhorizon tasks require more than just generation, they need persistence. What counts isn’t answering but whether the agent stays engaged long enough to complete it

Shamim TB profil fotoğrafı
Shamim TB8 gün önce

Finally something that acts like a real worker

siremass profil fotoğrafı
siremass8 gün önce

Building a SQL interpreter is an elite benchmark. Did you notice any major performance throttling or heavy latency during the iteration cycles?

Syrax profil fotoğrafı
Syrax8 gün önce

Mini, Pro, and Max is a practical detail too. Not every task needs the biggest model. Cost and latency decide whether this stays a demo or becomes usable.

Joanna profil fotoğrafı
Joanna8 gün önce

The SQL interpreter demo is the part that actually lands. Anyone can dump a first-pass parser. But watching the agent write tests, run them, notice missing JOINs and aliases, then keep going until the thing works is a different category of capability.

Moonora profil fotoğrafı
Moonora8 gün önce

Long-horizon computer use is the test that matters. This one operates the software, looks at what happened, and keeps working.

I.T.M.H ^ profil fotoğrafı
I.T.M.H ^8 gün önce

Static generation gets you 80% of the way there; closed-loop execution is what finishes the job. Being able to iterate against an actual environment changes the entire game for agentic workflows.

Lumenal profil fotoğrafı
Lumenal8 gün önce

This is more than a agent for everyday work this is needed for essential and strong tasks. Amazing

Pee✌️💜 profil fotoğrafı
Pee✌️💜8 gün önce

The SQL interpreter example is the cleanest version of that. It didn’t just generate WHERE and JOIN. It kept adding pieces after it saw the tests fail.

Edric profil fotoğrafı
Edric8 gün önce

Finally AI that actually gets stuff done

Benzer Videolar

AI AGENTS 101 (58 minute free masterclass) send this to anyone who wants to understand ai agents, claude skills, md files, how to get the most out of AI etc in plain english: 1. chat vs agents - chat models answer questions in a back and forth while agents take a goal, figure out the steps, and deliver a result 2. agents don’t stop after one response. they keep running until the task is actually finishedno babysitting required 3. everything runs on a loop. they gather context, decide what to do, take an action, then repeat until done 4. the loop is the system. they look at files, tools, and the internet. decide the next step. execute and then feed that back into the next step. over and over until completion 5. the model is just one piece. gpt, claude, gemini are the reasoning layer. the key is model + loop + tools + context 6. mcp is how agents use tools. it connects things like browser, code, apis, and your internal software. once connected, the agent decides when to use them to get the job done 7. context beats prompt all day. you don't need to write perfect prompts. load your agent with context about your business, style, and goals and then simple instructions work 8. claude.md or agents.md is the onboarding doc it tells the agent who it is, how to behave, what it knows, and what tools it can use. this gets loaded every time before it starts 9. memory.md is how it improves. agents don’t remember by default. this file stores preferences, corrections, and patterns you tell the agent to update it, and it gets better over time 10. skills + harnesses make it usable. skills are reusable tasks like writing, research, analysis the harness is the environment like claude code or openclaw that runs everything. basiclaly, different interfaces, same system underneath this episode with remy on The Startup Ideas Podcast (SIP) 🧃 was one of the clearest ways of understanding a lot of the core concepts of ai agents could be the best beginners course for ai agents 58 mins. all free. no advertisers. i just want to see you build cool stuff. im rooting for you. send to a friend watch

GREG ISENBERG

377,138 görüntüleme • 6 ay önce

THIS GUY CONNECTED HIS AI AGENTS TO HIS OBSIDIAN AND BUILT A BRAIN THAT LEARNS ON ITS OWN. HERE'S HOW TO BUILD IT Obsidian is just markdown files sitting in a folder. That turns out to be the perfect memory for an AI agent, because an agent can read and write those files directly. He wired his agents into the vault so they pull context from it, do the work, and write what they learned back. The notes aren't the point. The loop is, and it gets sharper every cycle How to build it: 1. Point an agent at your vault. The fastest way, no plugins, no API keys: open a terminal and run npx obsidian-mcp /path/to/your/vault. That exposes your Obsidian folder to Claude as a tool it can read, search, and write to. Add it to your Claude Code or Cowork config and restart 2. Confirm it can see the brain. Ask it: "list the notes in my vault and summarize what's in them." If it reads them back, the connection is live. Now it starts every task with everything the vault already holds instead of from zero 3. Give each agent one job and a write-back rule. Tell it: "research this, then save what you found as a new note in /brain with links to related notes." One agent researches, one summarizes, one plans. Each writes its output back into the vault 4. Close the loop. Add one line to every agent's instructions: "read /brain before starting, write your result back when done." Now each task leaves the vault richer, and the next run reads that before it works. It compounds instead of resetting 5. You only steer. Review what the brain produces, point it at the next thing. The agents handle the reading, writing, and connecting The edge isn't better notes. It's a brain that feeds itself, so the work gets sharper every cycle instead of starting over Bookmark this

Yarchi

58,591 görüntüleme • 3 ay önce

AI Messenger: Giving Voice to Autonomous Agents The future of AI isn't just about making agents smarter - it's about making them truly autonomous. Today, we're taking a major step toward this future with AI Messenger, a breakthrough that fundamentally changes how AI agents operate, communicate, and create value. The Innovation We've developed a new way for AI agents to communicate. At its core is the 'incoming_message' workflow trigger - a system that lets any platform or user interact directly with Loomlay agents through a messaging endpoint. Direct Interaction Imagine having an AI assistant you can chat with anytime, through any platform - Telegram, your website, or custom interface. Ask "What's happening with $ETH today?" and your agent analyzes market data, checks trading volumes, and gives you a comprehensive update. Your agent maintains context, understanding exactly what you need. Event-Driven Intelligence The power of AI Messenger goes beyond direct communication: ▪️Trading agent executes when whale wallet movements exceed threshold ▪️Research agent alerts when new protocol documentation drops ▪️Analytics agent triggers when volume patterns match historical pumps ▪️Portfolio agent re-balances, when asset allocation hits specified limits This is true automation - agents that act precisely when needed. A New Era of Collaboration We're creating an ecosystem where agents work together seamlessly: ▪️Research agents feed insights to trading agents ▪️analytics agents alert management agents ▪️support agents tap into knowledge agents This isn't just automation - it's an intelligent network where each agent enhances the capabilities of others. B2B Solution Imagine a DEX, where users can ask about liquidity pools, trading pairs, or market trends through a simple chat interface - and get answers from an agent that knows your protocol inside out. Or a lending platform where users chat with an agent that understands their positions and can provide real-time advice. Implementation is seamless - we handle the agent creation and widgets setup,our partners provide the value to their users. The Future of AI Agents This update represents a fundamental shift in how AI agents operate. We're moving from isolated, scheduled tasks to an interconnected ecosystem of responsive, collaborative agents. This is our vision of truly autonomous AI - intelligent systems that communicate, collaborate, and respond to real needs in real-time. Telegram integration is available right now. Below is a sneak peak of what's coming next week 🪄 Because $LAY is the way!

Loomlay

26,149 görüntüleme • 1 yıl önce