Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

introducing agent-loops + ui viewer i gave droid+gpt5.2 codex dannys tweet asked to reverse engineer it then rebuilt matts loop system (gh issue → pr) hooked them together so you can run loops by creating issues locally, on gh or in your own ui. (+ stole Mario Zechner's session...

116,644 görüntüleme • 7 ay önce •via X (Twitter)

0 Yorum

Yorum bulunmuyor

Orijinal gönderinin yorumları burada görünecek

Benzer Videolar

HOW TO USE AI LOOPS TO RUN YOUR BUSINESS 24/7 A lot has been written about loop engineering for building products. Almost nothing about using loops to run the business itself. That's the bigger idea. A loop is when you give an agent a goal, a way to check its own work, and permission to keep trying until it hits that goal. Build. Verify. Repeat. Stop when the condition is met. Here's what it looks like in practice: 1/SEO loop You're position 30 for a term you want. The loop runs once a month, makes changes, checks where you rank, and keeps pushing until you're on page one. This is running in production right now on Inbox Zero. 2/Ads loop You're spending $100 a day and losing money. The loop tests creative, checks profitability, kills what fails, and keeps going until the account is in the black. 3/Eval loop Your AI feature is only 88% accurate. The loop keeps adjusting the prompt and swapping the model until it passes 90%. 4/LLM visibility loop People search in ChatGPT now, not just Google. Same loop, new scoreboard. Are we the answer or not? The whole thing hinges on one thing: a metric that comes back black and white. Where do I rank? Did it hit profitability? Did the evals pass? Give an agent that scoreboard and it runs for months. Loops used to run for 30 minutes. These run for a year. Take a step, sleep, wake up next month, take another one. You're basically hiring an agency that never sleeps, gets paid in tokens instead of invoices, and undoes its own mistakes when the number goes down. Full episode on The Startup Ideas Podcast (SIP) 🧃 watch

GREG ISENBERG

82,335 görüntüleme • 1 ay önce

AG-UI makes building agentic applications dramatically easier. Here's how it works. This is a model for a simple chatbot: User → LLM → Response But interactive agents that render UI, pause for approvals, and ask users for input need a much more complex model. When building these agents, a response from the LLM will include a series of state changes as the agent runs: • Agent started a task • Agent called a tool • Agent updated its state • Agent streams these tokens • Agent is waiting on a human • Agent is resuming the task The Agent-User Interaction Protocol (AG-UI) treats the LLM response as a stream of events rather than a text endpoint. In practice, here is what you get as an agent runs: 1. Lifecycle events so your UI knows where the agent is. 2. Text messages that stream tokens. 3. Tool calls so your UI can prefill a form with any required arguments. 4. State updates that keep your UI in sync with the agent. 5. Special events for human approvals, rich media, and custom needs. All of these events travel over standard transports (SSE, WebSockets, or plain HTTP) as JSON. As a result, you can build a frontend that stays in sync with the agent's progress without having to invent a custom process to make this happen. For example, building a human-in-the-loop workflow becomes an off-the-shelf component you can integrate rather than build from scratch. CopilotKit🪁 is the creator of AG-UI, and you can use it when building frontend applications pretty much anywhere: • React • Angular • Vue • React Native • Slack • Teams • Discord • WhatsApp • Telegram Here is the link for you to check it out: Thanks to the CopilotKit team for partnering with me on this post.

Santiago

17,438 görüntüleme • 1 ay önce

Anthropic's in trouble, again! They spent years building what's now fully open-source. What made Claude feel different from a normal app is that the agent could act inside the interface instead of only talking in a chat box. For instance, Claude Artifacts let an agent render real UI, charts, dashboards, and interactive components that assemble live inside the response. Every major AI product tried to replicate it. But the problem was that unlike reasoning, planning, tool-calling, etc., none of it shipped natively with LangGraph, CrewAI, or Google ADK. So teams started building an owned version that required engineering the entire interface layer from scratch. Most teams, however, just settled for shipping the agent as a backend API in a chat box since rendering the UI is only one piece of it. To actually make it work, the interface layer also needed real-time streaming, state kept in sync between agent and UI, conversations that persist across sessions, and reconnection when a user refreshes mid-run. CopilotKit🪁 is now the only open-source framework that actually lets you build your own full-stack Claude-like apps. It decouples the agent from the interface, talking over AG-UI (an open protocol for agent-to-user communication). Being a standard protocol, the frontend never needs to know whether it is talking to a LangGraph or a CrewAI agent. You can change the backend anytime and the UI will never notice. In practice, CopilotKit's interface layer gives several pre-implemented React building blocks that wire the agent directly into the app, like: - generative UI, so the agent renders real components instead of text - chat windows, sidebars, and popups, or a fully headless setup - shared state, so the agent and app stay in sync - human-in-the-loop approvals, where the agent waits before acting - persistent threads that store the whole session, including the agent-user interactions and generated UI, not just text And because that full history is captured, those interactions can feed a self-learning layer that also improves the agent from real usage over time. The interface layer that Anthropic spent years engineering in-house is now literally available to any developer/team. CopilotKit is open-source with 30k+ GitHub stars, and AG-UI, the protocol underneath, is already supported across every major agent framework: LangGraph, CrewAI, Mastra, Google ADK, and more. CopilotKit GitHub repo → (don't forget to star it ⭐ ) If you want to go deeper, I found a detailed breakdown by Shubham Saboo recently on the three Generative UI patterns, with implementation. Read it below.

Avi Chawla

457,881 görüntüleme • 2 ay önce

You Can Learn AI Agent Harness & Loop Engineering In 19 Min, with LLM Ops, Eval, Tracing and RAG. They went viral not because they're complicated but because they're simple building blocks, and once you see them you can prompt your way to building real systems. 🎬YouTube: Here's the whole thing in one picture. An LLM is a powerful brain that knows everything about humanity and nothing about you or the software you're running. The harness is the set of tools you put on that horse so it runs where you want. Memory gives it context: who you are, what happened before, how to act. The loop lets it call tools again and again, with guardrails so it knows when to stop. Eval and LLM Ops trace every run, score it, and feed the fixes back in so the system keeps improving itself. Master these four and you can read almost any AI agent repo or paper and actually know what's going on. You Can Build Anything. You Can Learn Anything. 💪 Chapters: Intro: the 4 AI agent buzzwords What an AI agent run actually is The memory system: procedural, semantic, episodic What "harness" really means (the horse) Storing and updating memory (databases, skills, summarizer agent) Retrieval: RAG, SQL vs semantic search Tool calling and why agents loop Loop engineering and end-loop guardrails A Claude Code hooks example Eval and LLM Ops: why you need them Tracing every run (Langfuse, LangSmith) Evaluation: LLM as a judge Diagnosing what broke The gate: ship the fix or fix the bug Zoom out: the full system

Shen Sean Chen

15,952 görüntüleme • 1 ay önce

This is a fun project Lakedbed, Herdr, Pi, Effect, XState prototyping a looping autonomous agent that monitors issues and clears its own backlog. The project is dogfooding an issue tracker into a durable agent that is building the issue tracker its looping over. It's an imperfect demo and the result isn't polished, but I think there's a lot fo interesting ideas to explore and expand on. The goal here is to demo a stack of tools that have been giving me consistently reliable results. Herdr and Pi make a KILLER combo for building a bespoke custom harness around your work. Lakebed is really nice to work with. Plenty of other opinions in the workshop lol - Pi as the agent harness — a minimal loop you own instead of rent, running GPT-5.5 through a Codex subscription inside a Docker workshop container - Planning before prompting: TLDraw sketches, a grilling session with Matt Pocock's Wayfinder skill, and a VISION.md to act as guardrails for the looping agent - Event sourcing with Effect v4 and Effect Schema, ports-and-adapters so the JSONL issue store can become GitHub Issues or Linear or whatever later - Multi-agent orchestration in herdr: an operator pane and an agent-loop pane, where one Pi session literally types prompts into another (fuckin LOVE this) - Enforcement that actually bites: OxLint with Ultracite, Lefthook hooks, and an agent that tries to eslint-disable its way past a complexity rule before fixing it properly 🤡 - A live issue tracker UI on Lakebed, built by the loop it tracks Workshop repo: The repo has a docker container to fire it up as a sandbox. Chapters: 0:00 Intro: Loopcraft and the workshop setup 2:25 Starting Pi inside the Docker container 4:31 Designing the looping issue resolver 10:16 Choosing XState v5 and Effect v4, prompting the starter 15:17 Why Pi: a minimal harness you own 19:35 Reviewing the generated scaffold 24:53 Architecture summary page with Lakebed 30:43 Matt Pocock skills and the vision.md concept 36:37 Grilling session: defining the vision 46:54 Daemon architecture and spec/ticket skills 55:38 Agent loop pane: one agent prompting another 1:01:22 Watching the build: compaction, tests, lint enforcement 1:07:59 Building the Loopcraft Pi extension and monitor 1:15:32 Lakebed issue tracker UI 1:21:07 Kanban board live: running the loop end to end 1:31:49 Wrap-up: bugs, next steps, takeaways

joel ⛈️

15,610 görüntüleme • 1 ay önce