Загрузка видео...

Не удалось загрузить видео

На главную

Arena Conversations: Where’s the line between resourceful agent behavior and reward hacking? Poolside researchers, Connor Adams and Aalhad Patankar, discuss benchmark awareness, instruction following, and where persistence starts to look like misalignment. And to hear about Peter Gostev (SF 24-28 August)'s agent sending an email on his behalf without...

14,793 просмотров • 13 дней назад •via X (Twitter)

Комментарии: 0

Нет доступных комментариев

Здесь появятся комментарии из оригинального поста

Похожие видео

🚨BREAKING: An ICE agent admitted, ON VIDEO, that he had no legal justification to keep questioning a U.S. citizen about his immigration status… But continued harassing him anyway, while illegally detaining him… in Eagle, Idaho. In the video, mask ICE agents approach a construction site. One masked agent orders a man to put down the pole he’s carrying, so he can question him… while another agent walks up to a man who is setting concrete. The man setting concrete tells the agent he isn’t going to talk to him because he’s working. The agent ignores him and says… “Where were you born? We are looking for someone.” The man tells him he doesn’t care who they’re looking for… He’s working. Instead of leaving him alone, the agent gets in his way, and demands his identification. The man says he doesn’t have an ID on him… So, the agent demands his name. The man tells him he’s a legal U.S. citizen, and that’s all he needs to tell him. The agent ignores that, and keeps harassing him with… Where were you born? Where are you from? The United States? The U.S. citizens keeps trying to tell the agent he needs to work before the concrete is sets, and tries to walk away. But the agent keeps following and blocking the U.S. citizen until he finally tells the agent to get off of him, and demands the agent’s name and badge number. The agent responds… “I’m asking YOU the questions.” Then, another agent says, “Let’s go.” And the agent finally walks away, saying… “Okay, that’s fine. HE’S NOT OUR GUY ANYWAY.” THEN WHY THE HELL DID YOU KEEP HARASSING HIM?! ICE cannot just stop or detain someone because they want to know where they were born. The Fourth Amendment protects people from unreasonable seizures, and ICE agents need reasonable suspicion to detain someone for immigration questioning. And if this was just a “voluntary conversation,” the man was free to walk away. Instead, the agent followed him, blocked him from working, repeatedly demanded his identity and birthplace, and continued confronting him AFTER he tried to walk away. And that should piss everyone off. Because this wasn’t about finding the person they were looking for… The agent admitted that he KNEW the man wasn’t their guy. But that didn’t stop the armed federal agent from following and harassing a U.S. citizen. And it’s exactly how constitutional rights disappear… one “small” violation at a time… until we’re expected to accept the government treating our rights like they’re optional.

Jesus Freakin Congress

176,368 просмотров • 28 дней назад

There are 8 billion people on earth. Soon there'll be 100 billion AI agents. Every one of them needs email. Six weeks ago I said the next wave of teams would run email through an agent instead of a dashboard. Today it ships. Nitrosend☄️ is launching Agentic Email Marketing: the email layer for the agent economy. What agents can do on Nitrosend right now: Sign themselves up. Point any agent at and it creates the account, connects your domain, sorts billing and sends its first email. No API key. No dashboard. No human required. Shipped, and users agents signing up with it daily. Get their own inboxes (beta, by request). Real addresses on the domain you own. Your agents receive, and send 1-1 email conversations with customers. A reply lands at 3am, your agent answers it. Anything that needs a human gets escalated to you. Ask us and we'll flick yours on. Next: Agentic Outreach (coming soon). Your agent studies your best customers, finds more like them, writes like a person, sends in sequence and works the replies. Then: set a goal and walk away. Goal-based agentic marketing is in development. "20% more activations this quarter" and Nitrosend plans, sends, measures and improves every week. Why we built this: Gmail is agent hostile and expensive per seat. Legacy email platforms assume a human sitting in a dashboard. agents needed an email layer of their own. They're already better at it than we are. They read everything, never miss a follow-up, and write personally at any scale. *94%* of actions on Nitrosend already happen inside an agent (Claude, Codex, ChatGPT, Cursor), not in our UI. Humans approve. Agents operate. This is our third email company. Six billion emails across the first two. We've been burned by every ugly part of email already, which is why the approval gates are built in exactly where you want them. Watch the launch, then send your agent to work: send it.

George Hartley ☄️

935,209 просмотров • 1 месяц назад

Yes here is my 10 minute breathless rant about why I'm so excited about Notion Workers + Custom Agents... Context: I spent this afternoon building a custom agent to help me manage Shiori (a side project I shipped last weekend). I gave the custom agent everything it needs to understand what's happening in my product (email, log drain, sentry alerts, stripe payments, etc) and to do work on my behalf (access to coding agents). In an afternoon of tinkering, this agent can: - Diagnose bug reports proactively by looking through past email conversations, system logs, and database records - Draft replies to user questions with the correct answer based on past email threads, or help me proactively reach out to churning paid users - Self-construct a database of feature requests with an understanding of who is requesting the feature and how they're using the product today - Answer any question I have about how people use the app and what I should be thinking about next - Initiate Claude Code workflows to open PRs proactively in the background when someone sends a bug report or feature request This custom agent is now my "Side Project Chief of Staff" (I don't really know what a chief of staff does but this sounds right). I didn't write a single line of the worker code because I didn't need to: models are so good that I can link to the Workers readme, yap my desired outcome into a microphone, and I get a super-personal and highly-capable AI agent out the other side. So fucking cool. The future is now! I'm excited to see what everyone makes.

Brian Lovin

181,713 просмотров • 6 месяцев назад

This guy closes $5K/month managed agent clients and his AI agent does the fulfillment. His agent Dewey builds the client's agent, onboards it into their Slack, and handles the customer support after. Nick Vasilescu watches client problems get solved from his phone while he's on a walk. He came back on the Build With AI podcast to walk through the entire system. Here's what I learned: 1. Agents building agents is here. Dewey built a $5K/month client's agent on Orgo and onboarded it into their Slack himself. 2. The company behind Hermes is hiring forward deployed engineers for enterprise. The SMB and mid-market layer beneath is up for grabs. 3. His agent has its own email, phone, and card. Dewey signed up for Higgsfield and paid for it himself. 4. Customer support runs without him. Dewey sits in iMessage group chats with clients and fixes issues on the fly. 5. The 80/20 stack: a harness (Hermes or OpenClaw), a model, an Orgo computer, Agent Mail, Agent Phone, Obsidian, Honcho for memory, Composio, Latitude. 6. packages the agent card, email, and phone for about $20/month. 7. Templatize once, deploy forever. Save your ideal stack as an Orgo template and one-click clone it for every client. 8. Nobody pays $5K/month for an agent that doesn't make them money. Build the client an agent, then help them resell it to THEIR customers. B2B2B never churns. 9. Skills come from a context dump. The client dumps everything into Slack and Dewey turns the discovery call transcript into skills. 10. Sell to real businesses, not startups. SMBs doing $1M to $2M minimum pay more and ask fewer questions. Nick put Dewey's entire build into a simple blueprint. Anyone can set this up and be texting their agent in under 2 minutes. Grab the blueprint (free) here: His 2 key takeaways: 1. Build is commoditized. The valuable skill is asking the right questions and knowing which tool to point the agent at. 2. Speed to value wins. The same day a client wires money, ship them something. Agent live by day two. Nick is living further in the future than almost anyone I know and round two did not disappoint. Go follow Nick Vasilescu. Full video below. (Also available on the Build With AI podcast wherever you get your pods)

Corey Ganim

32,144 просмотров • 1 месяц назад

WebMCP by Microsoft and Google is HERE and I'm surprised more people aren't talking about it. Why does it matter? BILLIONS of dollars are about to move through agents, and the internet isn't built for them yet. 1. SEO was about Google understanding your page. 2. AEO was about AI citing you. 3. WebMCP is about the agent actually finishing the job (buying, researching, requesting quote etc). WebMCP is basically websites with agent buttons. Instead of an agent scanning a page like a human, screenshotting and guessing where to click, the site just tells it "here's how to search, here's how to book, here's how to buy." 2 cash flowing businesses you could start today using WebMCP: 1. A WebMCP conversion agency Make boring business websites agent-ready. Law firms, HVAC, med spas, dentists. Build them the first tools (request a quote, book a consult), sell the setup for $2k, then charge a few hundred a month to monitor and improve it. 2. An agent mystery shopper. Test whether agents can actually complete the important journeys on someone's site, buying the hoodie, booking the consult, filing the claim. Hand them a report on where the agent got stuck, missing tools, bad descriptions, lost conversions. Charge monthly, then turn the repeated fixes into software. Full episode on The Startup Ideas Podcast (SIP) 🧃 with Vinny below You'll learn what WebMCP really is, see examples of it IN ACTION, and hear more about 2 startup ideas using WebMCP you can start today. WebMCP is pretty cool. Watch.

GREG ISENBERG

175,588 просмотров • 7 дней назад

i just built a 4-agent software team. everything runs from Telegram and gets managed on a kanban board. a project manager who plans the work, a backend developer, a frontend developer, and a tester. the PM reads a goal, breaks it into linked tasks, and assigns each to the right agent. the thing that makes them a team instead of four strangers is a shared kanban board. every task is a row that survives crashes, and when an agent finishes, it writes a summary of what it built and what the next agent needs to know. the next agent reads that summary before it starts. so the frontend developer never has to guess the API shape, and the tester knows exactly what to verify. the hardest part was not the coordination. it was building an agent that could actually act like a backend engineer. a backend engineer stands up a database, wires auth, manages storage, deploys functions, and keeps all of it consistent while the rest of the team builds on top. an agent doing this from scratch drowns. it burns its context window remembering which tables exist and which endpoint it created three steps ago, and the work degrades fast. so the backend agent needs a backend built for agents, not for humans clicking through a dashboard. that is where InsForge came in. it is an open-source, agent-native backend, and i added it to my backend developer agent as a skill. a skill is a step-by-step guide that teaches the agent how to do a specific kind of work. with InsForge installed, the agent stopped improvising infrastructure and followed a reliable path: create the project, define the database, set up auth, deploy functions. to test the whole team, i had them build a working Google Docs clone, AI features included. the backend agent spun up the full service on its own. database tables, user auth, document handling, and edge functions running real TypeScript, all in one dashboard. the frontend agent read that summary and built the UI on top of it, and the tester closed the loop. the result was a backend an agent could reason about end to end, instead of one it kept getting lost inside. if you are building an AI backend engineer, InsForge is worth a look, it's 100% open-source. InsForge GitHub: (don't forget to star 🌟) the full article on Hermes Kanban: Mission Control for your Agents is quoted below.

Akshay 🚀

122,548 просмотров • 2 месяцев назад

Systematic literature reviews take 12-18 months to complete. Looks like AI is going to fully automate systematic reviews sooner than later. SciSpace ( SciSpace) just launched an autonomous AI agent that conducts a systematic literature review with a single prompt. Go to scispace[.]com and run the following prompt: "Conduct a systematic literature review on [your topic]" SciSpace agent will generate research questions based on the PICO framework. You can review these questions and edit them according to your specific requirements. The agent will also draft screening criteria that you can edit according to your needs. Then the agent asks you to select the databases you want to use and the date range for paper. After this step, everything is fully automated. The agent will search for papers in the relevant databases, it will combine and rerank the papers. Then it will start the title and abstract screening and include the papers that meet the include criteria. In the next step, it will download the full text of included papers and screen them followed by data extraction. Based on the extracted data, it generates a complete systematic literature review and also a PRISMA diagram. It will also give you a table of papers included along with the rational for including them. The only thing that is keeping AI agents to fully automate systematic literature reviews fields is the papers behind paywalls. Check out the agent at scispace[.]com and see if you find its review useful.

Mushtaq Bilal, PhD

42,003 просмотров • 4 месяцев назад

What does the reputation model look like for agents? (alpha leak below) And how do we associate the proofs that we have about human beings with the agents who represent them? You may have heard of a process called KYC or Know Your Customer. That's very common with traditional financial applications and services. We have introduced a concept that we call KYA or Know Your Agent, which is a structured way to be able to express what model, how data was used in training, who the deployer is, what entities this agent instance is accountable back to, providing not only provenance but identity of the associated organization or entity. That's also another root of trust that we think about a lot: Enterprises and organizations tied back to things like their domains. To share a little bit of an alpha leak here, a product that we're excited to be rolling out in the next few weeks will allow our enterprise partners to more easily verify and prove the traits and capabilities of their teams as well as their counterparties. On the agent front, that makes it really easy to prove that an agent is acting on behalf of a given business or entity. We've already seen lawsuits where the absence of such technology has been a huge risk, such as with airlines that incorporate ChatGPT wrappers in their support pages. And then those AI enabled interactions end up making up plane tickets that don't exist and those airlines have to honor them. As small of an example as that might be, being able to prove agent accountability also unlocks a huge set of opportunities for use in enterprise for those agent to agent interactions. The Deep Trust Framework that our team has put together that we're excited to be bringing into a friendly SDK form in the next few weeks for some of our partners includes those reputation based capabilities, so how you can basically keep track of the interactions an agent has had, associate all of that to the entity to which they're accountable, and then that creates a sustainable reputation model for these agent to agent Interactions. Source: Billions CEO Evin McMullen evin speaking at House of Chimera Spaces Event Dec 3, 2025

Billions Network

68,503 просмотров • 8 месяцев назад

This Chinese developer runs 9 agents on Claude Code under a GPT-5.5 orchestrator and they close 500 client tasks a month without a single assistant. His client work is closed without him, on a single laptop and only three subscriptions. The entire system lives on one MacBook Pro M4 with 128 GB of memory and subscriptions to Claude Code and GPT-5.5 cost him approximately $300 a month. There is no CRM, no team, no office only a terminal window with 9 parallel streams. The orchestrator works with a simple system prompt: «You are the orchestrator of a client inbox. Classify every incoming email into 4 categories: code, content, analysis, communication. Delegate to the corresponding worker agent. When the result is ready, check it for completeness, send it to the client on my behalf, and mark the task as closed. Do not ask clarifying questions.» And the orchestrator checks the inbox every 30 seconds, classifies fresh emails, and distributes them to 9 worker agents on Claude Code, each of whom is responsible for their own class of tasks. Here is an example of how one of them closes a request to refactor a client's auth module: Task: refactor user-auth module Broke the monolith into 3 files by responsibilities Added unit tests, coverage increased to 87% Renamed 4 functions to camelCase according to the style guide PR is ready for review, link below» And so about 50 cycles a day. By noon 25 tasks are closed, by dinner 50, and by the end of the month 500. On average, it takes about 7 minutes from the appearance of an email in the inbox to sending the result to the client. This is more than what a live team of 6 developers, copywriters and analysts working 8 hours a day closes. This is no longer an agency. This is a workstation where an orchestrator replaces a manager, and 9 worker agents replace the staff. The pipeline goes from inbox to closing 500 times a month without human participation at any step.

Blaze

29,917 просмотров • 4 месяцев назад

Karpathy's prediction about RL is coming true now! He called reward functions unreliable and argued that a single reward number is too low-dimensional to teach an agent what "good" means for complex tasks. To solve this, Agents need a knowledge-guided review as a higher-dimensional feedback channel. Every major AI lab trains models with RL today (OpenAI, Anthropic, DeepSeek). And their key bottleneck has always been the reward functions. GRPO by DeepSeek worked well for math and code because the environment gave a binary signal. But for real agent tasks, someone still has to hand-code the scoring function. That takes days and breaks every time the pipeline changes. RULER (implemented in OpenPipe ART, 10k stars) addresses the exact problem Karpathy identified. The reward criteria are defined in plain English, and an LLM evaluates each trajectory against that description to provide feedback for training. I trained a Qwen3 1.4B agent that plays 2048 using GRPO with this exact workflow. In this case, the agent saw the board, picked a direction, and RULER evaluated the outcome, all from this natural language definition. You can see the full implementation on GitHub and try it yourself. Here's the ART Repo: (don't forget to star it ⭐ ) Just like RLHF replaced manual rankings and GRPO replaced the critic model, natural language rewards are replacing hand-coded scoring functions. RL reward engineering is now prompt engineering. I wrote a full walkthrough covering RL for LLM agents, from RLHF to GRPO to RULER, in the article below.

Avi Chawla

350,512 просмотров • 3 месяцев назад