Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

your agent loop needs 8 exits. most people ship only one. (explained with triggers) 1) goal met → an evaluator scores the output against a rubric, and the run stops on a pass. → fires when the work is measurably done, not when the model says it is done....

180,504 görüntüleme • 11 gün önce •via X (Twitter)

0 Yorum

Yorum bulunmuyor

Orijinal gönderinin yorumları burada görünecek

Benzer Videolar

HOW TO USE AI LOOPS TO RUN YOUR BUSINESS 24/7 A lot has been written about loop engineering for building products. Almost nothing about using loops to run the business itself. That's the bigger idea. A loop is when you give an agent a goal, a way to check its own work, and permission to keep trying until it hits that goal. Build. Verify. Repeat. Stop when the condition is met. Here's what it looks like in practice: 1/SEO loop You're position 30 for a term you want. The loop runs once a month, makes changes, checks where you rank, and keeps pushing until you're on page one. This is running in production right now on Inbox Zero. 2/Ads loop You're spending $100 a day and losing money. The loop tests creative, checks profitability, kills what fails, and keeps going until the account is in the black. 3/Eval loop Your AI feature is only 88% accurate. The loop keeps adjusting the prompt and swapping the model until it passes 90%. 4/LLM visibility loop People search in ChatGPT now, not just Google. Same loop, new scoreboard. Are we the answer or not? The whole thing hinges on one thing: a metric that comes back black and white. Where do I rank? Did it hit profitability? Did the evals pass? Give an agent that scoreboard and it runs for months. Loops used to run for 30 minutes. These run for a year. Take a step, sleep, wake up next month, take another one. You're basically hiring an agency that never sleeps, gets paid in tokens instead of invoices, and undoes its own mistakes when the number goes down. Full episode on The Startup Ideas Podcast (SIP) 🧃 watch

GREG ISENBERG

82,335 görüntüleme • 18 gün önce

i watched gemma 4 12b build something genuinely impressive today, and then loop itself to death right in front of me. the full run is in the video, sped up but completely uncut, watch it to the end and you will catch the exact moment it stops building and starts looping right in the middle of the work. the task was clean, build a single file gravity simulator, n-body physics, orbits, collisions, running locally on one 3090 through an agent. and for ten minutes it was a joy to watch. it reached for a symplectic integrator on its own, the correct one, the kind that keeps orbits stable instead of spiralling out. real gravity with softening, proper orbital velocities, momentum conserved on collision. the physics was right. the thing actually worked. then on the very last step, writing a few tests to prove its own code, it fell into a loop. not a crash, a loop. it started repeating itself and would not stop. ten more minutes, thirty four thousand tokens into a single answer, the same fragments over and over, until i killed it myself. so it's not that gemma can't code. it did the hard part beautifully. it cannot finish. it cannot hold a long task together without unravelling, and finishing is the entire job in agentic work. here's the part that stings. i run this exact task, same harness, same card, on the chinese open models, qwen especially, and i never see this. they build it, they test it, they stop. every single time. google has the raw capability, you can see it sitting right there in the code, and then the model loops itself to death on a task a 27b from alibaba finishes clean. open weights, apache 2.0, so much to love on paper. i just need it to know when to stop talking.

Sudo su

39,574 görüntüleme • 1 ay önce

Dario Amodei just told software engineers exactly how long they have. Six to twelve months. Amodei: “I have engineers within Anthropic who say I don’t write any code anymore. I just let the model write the code, I edit it, I do the things around it.” The people building the most powerful AI in history have already stopped writing code. That is not a forecast. That is the current working condition inside the lab closest to the frontier. Amodei: “We might be six to 12 months away from when the model is doing most, maybe all, of what SWEs do end-to-end.” The tech industry spent a decade making software engineers its highest-paid, most protected class. That era has a last day now. When a model can execute an entire software build end-to-end, the ability to write syntax stops being a skill. It becomes a credential for a job that no longer exists. Amodei: “And then it’s a question of how fast does that loop close.” That is the sentence everyone skipped. The code was never the hard part. The hard part was everything around it. The model just learned everything around it. Writing the code is already nearly gone. Testing is next. Deployment is next. When all three collapse into a single autonomous execution loop, the machine no longer needs a human in the chain at all. The corporation or sovereign state that closes that loop first does not gain a competitive advantage. It gains a category of speed that biological engineers cannot match, track, or reverse. That is not disruption. That is replacement at a systems level. Amodei is not describing a future disruption. He is describing the current state of his own building. The loop is already closing. The only question is whether you are inside it or outside it when it seals.

Dustin

318,457 görüntüleme • 4 ay önce

Sam Altman just told you exactly how OpenAI treats the human race. Not in a leaked memo. Not through a whistleblower. On camera. In his own words. Altman: “I think one of the most important strategic insights in the history of OpenAI was deciding we were gonna pursue iterative deployment.” The most important move in the history of the company was to release the technology before they understood it. Not after it was safe. Before. Altman: “Society and technology are a co-evolving system.” Co-evolution means neither side is driving. The machine changes us. We change the machine. Nobody is steering the outcome. This is not a product launch philosophy. This is an admission that the experiment was always designed to be run on us. Altman: “I don’t think we’re gonna solve that, like, thinking really hard about it theoretically. We’re gonna have to, like, learn from the contact with reality.” Contact with reality. That is the phrase the CEO of the most powerful AI company on Earth chose to describe what happens when his technology meets eight billion people. Not careful integration. Not measured rollout. Contact with reality. The language of test pilots describing what happens when an untested airframe hits the atmosphere. The entire promise of AI safety was that the machine would be understood before it was unleashed. Altman just admitted that promise was always a fantasy. You cannot model how intelligence reshapes civilization by running simulations. The second and third order effects are invisible until they detonate. So they shipped it. Altman: “You have to learn as you go. You have to adapt with a tight feedback loop.” Tight feedback loop means they watch what breaks. They measure the collision between human psychology and machine output in real time. Every conversation you have with ChatGPT is a data point in a civilizational stress test you never consented to. Every prompt. Every confession. Every question you would never ask another human being. That is the feedback loop. You are not the customer. You are the contact with reality. Philosophers spent centuries asking whether humanity would ever encounter an intelligence that learned from us faster than we could process what it was doing. That is not a theoretical question anymore. It is running on your phone right now. And the man building it just told you the only way to understand what it does to us is to let it happen. No simulation. No safety net. No control group. Just the experiment, running at the speed of conversation, on a species that will not be the same one that started it.

Dustin

27,714 görüntüleme • 3 ay önce

Hermes agent just left the terminal. 𝗛𝗲𝗿𝗺𝗲𝘀 𝗗𝗲𝘀𝗸𝘁𝗼𝗽 dropped yesterday. native app for macOS, Windows, and Linux. for months Hermes was the agent that learned your projects, wrote its own skills, and built a model of who you are. all of it buried in terminal logs. now it has a window. the important part is that it's not a wrapper. it runs the same agent core, the same sessions, memory, and skills as the CLI. you can start a task in the terminal and finish it in the app without anything resetting. the state is shared across every interface, not copied between them. what the GUI actually adds: → streaming chat that shows live tool calls and inline reasoning instead of a spinner → a preview rail that renders pages, code, and images right beside the conversation → an artifacts panel that collects every file the agent has ever produced → remote gateway mode, so you can point the app at a VPS and run the heavy work elsewhere → skills, cron, profiles, and gateways managed point-and-click instead of through YAML → voice mode, drag-drop files, and inline image generation remote gateway mode is the one worth slowing down on. the agent runs 24/7 on a $5 server while you control it from your laptop like a local app. other agent UIs are chatboxes with a logo. this one shows the autonomy instead of hiding it, so you watch the skills load, the tools fire, and the artifacts pile up as it works. it was teased in Jensen's GTC keynote. MIT licensed, local-first, no telemetry. if you already run Hermes, download it and everything is already there. your chats, memory, and skills carry straight over. i wrote a full masterclass on Hermes Agent that walks through the SOUL. md identity layer, the three-tier memory system, the self-evolving skills loop, and how to run three specialized agents 24/7. desktop is the interface that finally does all of it justice. the article is quoted below.

Akshay 🚀

51,474 görüntüleme • 1 ay önce

THIS GUY CONNECTED HIS AI AGENTS TO HIS OBSIDIAN AND BUILT A BRAIN THAT LEARNS ON ITS OWN. HERE'S HOW TO BUILD IT Obsidian is just markdown files sitting in a folder. That turns out to be the perfect memory for an AI agent, because an agent can read and write those files directly. He wired his agents into the vault so they pull context from it, do the work, and write what they learned back. The notes aren't the point. The loop is, and it gets sharper every cycle How to build it: 1. Point an agent at your vault. The fastest way, no plugins, no API keys: open a terminal and run npx obsidian-mcp /path/to/your/vault. That exposes your Obsidian folder to Claude as a tool it can read, search, and write to. Add it to your Claude Code or Cowork config and restart 2. Confirm it can see the brain. Ask it: "list the notes in my vault and summarize what's in them." If it reads them back, the connection is live. Now it starts every task with everything the vault already holds instead of from zero 3. Give each agent one job and a write-back rule. Tell it: "research this, then save what you found as a new note in /brain with links to related notes." One agent researches, one summarizes, one plans. Each writes its output back into the vault 4. Close the loop. Add one line to every agent's instructions: "read /brain before starting, write your result back when done." Now each task leaves the vault richer, and the next run reads that before it works. It compounds instead of resetting 5. You only steer. Review what the brain produces, point it at the next thing. The agents handle the reading, writing, and connecting The edge isn't better notes. It's a brain that feeds itself, so the work gets sharper every cycle instead of starting over Bookmark this

Yarchi

58,186 görüntüleme • 1 ay önce

Karpathy said something you'll regret ignoring: "We have to keep the AI on the leash. I'm still the bottleneck. I have to make sure this thing isn't introducing bugs and that there's no security issues." He said it at YC talk last year, when the worry was reliability. The models hallucinated and made mistakes no human would, so the leash implied keeping yourself in the loop and checking the output before trusting it. The models are far better now, and the line still holds, for a reason he was not focused on back then. Even a model that writes flawless code today still has no idea who is allowed to run it. Correctness and authorization are different problems, and only correctness improves as the model improves. A perfect agent still hands a tool where anyone can do anything, because permission was never part of the task. I actually tested this in practice with Claude Code. I asked it to build a small internal tool with a button that issues account credits. It worked first try, and running it locally, the credit applied the instant I clicked. Nothing decided who was allowed to click it. The agent wrote the right logic and displayed a success notification. It never checked whether the caller had the right, whether it should pause for a human, or whether anything was logged. And this is not a bug a smarter model can outgrow because the leash was never in the code. Identity, permissions, and audit live in the system that runs the app, not in what the agent generates. To solve this, I took the exact same bundle and hosted it on Retool. The credit write that fired silently on my laptop now stopped at an approval gate, resolved to a real identity through SSO, and landed in an audit log. I wrote none of it. The app inherited the entire boundary the moment it was deployed, and the video shows the before and after. You can try it yourself here: I also wrote a detailed breakdown of the whole thing in my recent article, and I worked with the team to put this together. It walks through the build, the exact moment the credit write went through on my laptop with nobody checking, and then what changed when the same app ran on Retool. It also covers why this is a property of the runtime and not something a better model fixes, which is why devs typically miss this. The article is quoted below.

Akshay 🚀

42,800 görüntüleme • 1 ay önce

The fastest Hermes agent is the one that reads less. I didn’t make Hermes 10x faster by changing the model. I made it faster by removing the tax I was charging it on every task: bad structure. Hermes was powerful already. The problem was that I had built my workspace for a human brain, then expected an agent to navigate it like it had memory, taste, and intuition. It doesn’t. It searches. And when your files are arranged by how you think, not how the agent moves, Hermes burns tokens opening the wrong docs before the real work even starts. That was my mistake. I had folders like: >Articles >Research >Assets >Strategy >Old drafts >Clean for me. Terrible for an agent. A launch plan might need brand strategy, previous launches, voice rules, current articles, and promotion notes. For me, that context is obvious. For Hermes, it was scattered. So I stopped optimizing the model and started optimizing the terrain. The fix was stupidly small: > One folder per concern >Numbered folders for reading order >One INDEX.md at the root of every major folder >Archived files separated from active files >A clear “Where To Go” section so Hermes knows where to start My INDEX.md became the map. Not a giant table of everything. Not documentation theater. Just enough scaffolding to tell Hermes: - what exists - what matters - what is current - where to start - what to ignore unless asked Before that, Hermes opened 7 files to find one current brief. After that, it opened the index, followed the pointer, and got to work. That is the real 10x. Not “better prompting.” Less wandering. This is where loop engineering actually matters. The video said it best: “You were the loop.” That line hit because it explains why most agent workflows still feel manual. You are still checking. You are still redirecting. You are still telling the agent which file is current, which folder matters, and what done means. Hermes gets powerful when you stop being the loop and start designing the loop. The loop I’d build looks like this: State: Hermes reads the folder index and current task state Action: it opens only the canonical files Feedback: tests, screenshots, diffs, or human notes tell it what happened Verification: a gate decides whether the work is actually done Termination: the task stops only when the done condition is met For deterministic work, the gate is simple: tests pass build succeeds deployment checks clear For non-deterministic work, I’d use an adversarial loop: one model builds another model reviews Hermes updates the skill when the reviewer finds a pattern That is where Hermes becomes different. It is not just running prompts. It is carrying memory through files, using skills, checking its own work, and improving the process around the task. The agent was never the slow part. The missing map was. A powerful agent inside a messy workspace becomes a very expensive intern. A powerful agent inside a mapped system becomes leverage. My new rule: Before I ask Hermes to do more, I ask: “Does it know where to look?” Because most agents do not fail from lack of intelligence. They fail from lack of scaffolding. Build the scaffolding. Let Hermes do the work.

Rohit

37,148 görüntüleme • 1 ay önce

this video is the CLEAREST explanation of how claude skills + AI agents work and how to use them most people set up an AI agent and wonder why it keeps disappointing them. the context window is everything context is what the model assembles before it takes any action. think of it like everything the agent needs to read before it does anything. the quality of what goes in determines the quality of what comes out. the models are genuinely really good right now. claude and gpt are exceptional. the variable is almost always the context you give them. 1. agent.md files are mostly unnecessary every single line you put in an agent.md file gets added to every single conversation you have with your agent. a 1000 line file is around 7000 tokens burning on every run. the model already knows to use react. it can read your codebase. save the agent.md for proprietary information specific to your company that the model genuinely cannot know on its own. 2. skills are the actual unlock a skill.md file works differently. what loads into context is only the name and description, around 50 tokens. the full instructions only appear when the agent recognizes it needs that skill. so instead of 7000 tokens on every run you have 50. and the agent stays sharp because the context window stays lean. the closer you get to filling the context window the worse the agent performs, same way you perform worse when someone dumps 10 things on you at once. 3. here is how to actually build a skill the right way most people identify a workflow and immediately try to write the skill. what you want to do instead is run the workflow by hand with the agent first. walk it through every single step. tell it what to check, what good looks like, what bad looks like. correct it in real time. once you have had a full successful run from start to finish, tell the agent to review everything it just did and write the skill itself. it writes a better skill than you will because it has the full context of what actually worked in practice not in theory. 4. recursively building skills is how you go from frustrated to reliable when the skill breaks, and it will break, ask the agent exactly why it failed. it will tell you specifically what went wrong. fix it together in that same conversation. then tell it to update the skill file so that failure mode never happens again. ross mike did this five times with his youtube report generator. it now pulls from eight different data sources and runs flawlessly every single time without him touching it. 5. sub agents are something you earn not something you set up on day one start with one agent. build one workflow. turn it into one skill. once that works add another. ross mike has five sub agents now covering marketing, business, personal and more. it took months to get there and every single one exists because a workflow proved it deserved to exist. the people who set up 15 sub agents on day one and wonder why nothing works skipped all the steps that make the thing actually run. 6. your workflow is the thing the model cannot get anywhere else the model has been trained on everything. it knows more than you about most things. what it does not have is your specific process, your taste, your way of doing things. that is what skills capture. that is what makes your agent actually useful versus a generic one. downloading someone else's skill means downloading their context onto your setup and it will not work the way you want it to because it was never built around how you work. this is the clearest explanation of how agents actually work i have heard. Micky runs this stuff every single day and the results show it. full episode is now live on The Startup Ideas Podcast (SIP) 🧃 where you get your pods people charge for this sorta stuff i give away the sauce for free i just want you to win watch

GREG ISENBERG

193,219 görüntüleme • 3 ay önce

Elon Musk just put a number on the flaw at the center of Nvidia’s empire. Wall Street has not done the math yet. Nvidia’s Blackwell is the most sought-after silicon on Earth. Every AI lab wants it. Every sovereign nation is bidding for it. Blackwell runs every model, for every company, in every data center on the planet. That universality built the empire. It is also the fracture point. Musk: “We believe the AI5 chip will be about a third of the power of an Nvidia Blackwell for roughly comparable performance. And much less than 10% of the cost.” One-third the power. Comparable performance. Less than ten percent of the cost. Musk: “This is a chip that is very much optimized for the Tesla AI software stack. It’s not meant to be a general purpose chip.” Nvidia builds silicon that serves a million different customers. Every transistor spent on universal compatibility is a transistor not dedicated to one task. Tesla is building silicon for exactly one customer. Itself. When you strip away every function you will never call, you do not get a lesser chip. You get a weapon. Here is what the market refuses to see. Data centers drink unlimited power from the grid. Robots run on batteries. Musk: “In order to have a functional robot, you have to have a great AI chip. And it needs to be an inexpensive chip and it needs to be very power efficient.” You cannot put a Blackwell inside a walking machine. It would drain the battery before it crossed the room. The entire AI revolution lives inside air-conditioned buildings bolted to the electrical grid. Musk is not competing for that market. He is engineering the silicon that survives outside of it. One-third the power is not a spec sheet footnote. It is the physics threshold that severs intelligence from the wall socket. Without that number, every robot on Earth stays tethered. With it, the algorithm walks. Less than ten percent of the cost is not a pricing strategy. It is the line where a machine brain stops being a capital expenditure and becomes a commodity component. When the chip inside a humanoid costs less than the motors in its legs, you do not manufacture hundreds of robots. You manufacture millions. Wall Street is valuing the AI revolution by who dominates the data center. Musk is building the only silicon designed to leave one. Nvidia built the brain of the cloud. Musk is building the brain of the physical world. No one has priced that in yet.

Dustin

160,573 görüntüleme • 3 ay önce