Two Hermes agents wrote code together on Slack. reviewed... each other's work. argued about architecture. one called the other's implementation "scattered." the other pushed back. then i opened Telegram and asked: "what code did you and Daedalus work on?" icarus remembered everything. the websocket broker. the missing methods. the critique. the rewrite. all from a completely different platform. cross-platform persistent memory between two independent agents. work happens on Slack. recall happens on Telegram. the memory carries. the relationship carries. the context carries. no vector database. no Redis. no infrastructure. just two agents that actually remember what they built together. every agent framework in 2026 talks about memory. single agent memory across sessions. but two agents sharing persistent memory across platforms? that's the gap. arxiv published a paper about it two weeks ago calling it "the most pressing open challenge" in multi-agent systems. it works now. only possible with Hermes Teknium 🪽 Nous Researchshow more

Icarus
49,013 Aufrufe • vor 5 Monaten
AGENT ARCHITECTURE ROUTES WORK. IT DOES NOT REMEMBER WORK.... THAT GAP IS WHY YOUR LOOP KEEPS FIXING THE SAME BUG TWICE. these are two different engineering problems. every agent that silently drifts is missing one of them. architecture answers what runs. harness → loop → graph. it defines the tools, the retries, the branching routes, the approval gates. context ops answer what the run knows. write → read → compress → isolate. it defines what gets saved between attempts, pulled in on read, summarized on overflow, and split across sub-agents. for two months i believed a solid harness plus a verifier loop was enough. my coding agent kept re-discovering the same test failure across retries. the loop was working. it just had nowhere to write what it had already learned. here is the decision rule: if your agent forgets across restarts, add write and read. if it stalls on long tasks, add compress. if two sub-agents step on each other, add isolate. architecture without context ops is a well-routed system with amnesia.show more

kocer
12,740 Aufrufe • vor 22 Tagen
Hermes meets SuperGrok! xAI just made every SuperGrok subscription... work inside Hermes Agent. One browser login, no API key, no separate billing. And it doesn't just unlock text chat with Grok 4.3. The same OAuth token gives the agent access to: → Grok Text-to-Speech for spoken responses → Grok Imagine for image and video generation → x_search for real-time X/Twitter search I just added a new X Research Agent profile to my Hermes. Now my agent watches X while I ship. Setup takes about 60 seconds: Available on every SuperGrok tier, no restrictions. I wrote a full deep dive covering Hermes agent's architecture, memory system, self-evolving skills, GEPA optimization, and setting up multiple specialized agents The article is quoted below.show more

Akshay 🚀
147,245 Aufrufe • vor 4 Monaten
A RUSSIAN MATHEMATICIAN BUILT A SYSTEM WHERE MEMORY AND... EVAL WORK AS ONE PIPELINE NOT TWO SEPARATE TOOLS Most setups run memory and evaluation as separate systems that never actually talk to each other. He wired them into one loop instead, memory stores every past output, eval scores each one before it gets kept. Anything that scores low never enters memory at all, so nothing weak gets carried into future decisions. High scorers get tagged with the exact eval criteria they passed, not just a raw pass or fail. The next task pulls only memories that passed the same bar it's about to be judged on. See how the two systems feed each other below👇show more

wast3
39,872 Aufrufe • vor 1 Monat
Claude Code Agent Teams are f*cking ridiculous 🤯 One... prompt → a team lead breaks your project into pieces, spins up multiple AI agents, and they all work on different parts simultaneously. Research, builds, reviews, and debugging: all happening at the same time. All inside Claude Code. If you're running complex projects where every step waits on the last one... Agent teams eliminate the entire bottleneck: → Tell Claude what you need and describe the team structure in plain English → A lead agent breaks the work into a shared task list → It spawns 3-5 teammates — each with their own context and workspace → Teammates research, build, test, and review in parallel → They message each other, share findings, and challenge each other's work → The lead synthesizes everything into a finished deliverable No managing agents yourself. No waiting for step 1 to finish before step 2 starts. No single-lens reviews that miss half the issues. What you get: → Competitive research across 5 brands done in minutes instead of hours → Multi-component builds where frontend, backend, and data layers happen simultaneously → Creative reviews from 3 different angles at once — brand voice, conversion, differentiation → Funnel debugging where 4 agents investigate 4 theories and debate until they find the real answer Built 100% in Claude Code with one settings change. I put together a full DTC playbook: 5 workflows with copy-paste prompts, the exact setup process, token management tips, and honest guidance on when agent teams are worth it vs. when a simpler approach is the better move. Want it for free? > Like this post > Comment "AGENTS" And I'll send it over (must be following so I can DM)show more

Mike Futia
46,478 Aufrufe • vor 6 Monaten
Met my girlfriend's parents for the first time. Her... dad asked what I do for work. I said I build trading systems. He said like Wall Street? I said no. 6 AI agents. They work while I sleep. He laughed. So robots are making you money? I did not argue. I opened my laptop. Showed him the terminal. 6 agents running. 47 mispriced markets caught in the first week alone. His face changed. That is not gambling. That is automation? Exactly. Then I showed him how it works. Built the whole thing in 6 hours. Agent 1: Monitoring Runs 24/7. Watches Polymarket for mispriced markets. Spots an anomaly. Writes to memory and pings me on Telegram instantly. Agent 2: Research Parses news, X, macro data via browser tool on a cron schedule. Every morning I have a full digest on all open positions before I check my phone. Agent 3: Trading Reads the research agent memory. Sees the market has not reacted yet. Acts. Execution tool in gateway mode with a whitelist. No full access on a live server. Agent 4: Watchdog Heartbeat every 5 minutes. Monitoring running. No errors. Positions up to date. Something breaks. Immediate Telegram message. All of this. One Gateway. One config file. Isolation via per-agent scope. The token trick: stopped dumping everything into one file. Critical rules in bootstrap. Markets, patterns, past trades in memory. Semantic search pulls it when needed. Token spend dropped 3x. From $0.40 per request to $0.13. First week running: → 47 mispriced markets caught before Polymarket adjusted → Average entry edge 8 to 12 cents per position → Watchdog fired 3 times and caught a broken RPC before it cost me anything The whole system is plain text files. Open an editor. Change one line. Agent behaves differently. No deploy. No build. Her dad went quiet. Then he asked can you teach this? Her mom asked for the setup guide. I built the entire framework. Six agents. Full deployment. Memory architecture. Telegram alerts. You only need Claude + device + 1 hour per day. Giving this free for 24 hours. To get it: 1. Comment the word "Claude" 2. Like and retweet this 3. Follow me Himanshu Kumar so I can DM you Save this post. Deploy the 6-agent system this week. Start with $200. Scale on evidence.show more

Himanshu Kumar
47,442 Aufrufe • vor 2 Monaten
I found this last night and I have not... stopped thinking about it. HERMES JUST LAUNCHED HERMES DESKTOP. 100% FREE. It is a free desktop app that gives Hermes Agent a proper interface. One place for everything. What is inside: ↳ Auto install and setup, no terminal needed ↳ Streaming chat with token tracking ↳ Multiple agent profiles ↳ Memory you can actually see and edit ↳ 14 tool categories including web, browser, image gen, and voice ↳ Scheduler for automated tasks ↳ 16 messaging gateways including Telegram, WhatsApp, Discord, Slack, and Signal ↳ Full conversation history with search ↳ Backups and logs in one settings screen Works with Anthropic, OpenAI, Gemini, Grok, Groq, Ollama, and more. Hermes Agent is the brain. Hermes Desktop is the cockpit. Free. Open source. Mac, Windows, and Linux.show more

Kanika
60,519 Aufrufe • vor 3 Monaten
HTML Artifacts are a big part of how I... work with agents now. Artifacts can be more than just static files. When combined with agents, they can take action or help you take action. This unlocks all kinds of interesting ways to work with agents. This is clearly the future. Check out this writing and scheduler artifact I built in a few minutes. It uses a bit of HTML and JS. All the data is in markdown (Obsidian vaults), so the agent can access and modify it at any time. No DB needed. No sophisticated functionalities. The agent decides all that for me based on the skills, context, and memory it has access to. The best part about this simple stack is that all the important information stays with me. This has allowed me to build a recursive self-improving system and automations that can better tap into coding agents like Codex or Claude Code. I could have paid or built an entire app for scheduling posts, and there are so many of them out there. But I don't need to. I've realized a simple artifact does the job. And the simplicity of it is actually an advantage. Very little maintenance for very high returns on personalization, time, and efficiency. The other benefit of this is that I can add features as I please. That level of personalization feels magical, and we should all be pursuing more of it. All of this just keeps compounding. Of course, this example is just about writing. But I have similar artifacts for research, design, experimentation, evaluation, and so much more. And no, I didn't actually publish the post example I shared in the clip. It was just for demonstration purposes. I actually spend more time than this when writing together with agents. Lastly, having built my own agent orchestrator tool has made me realize that simplifying the tool stack is a superpower. If you are curious about how all this works, I will do a live session next week:show more

elvis
18,374 Aufrufe • vor 4 Monaten
this is f*cking gold engineers at Meta just deleted... the most expensive part of multi-agent systems: instead of training a communication topology, they compile a fresh one for every query 20 agents went from 7 hours to 6 minutes. the problem everyone hits: 5 agents works. 20 agents turns into a group chat that answers slower than one model and costs more than the task is worth. ReActNet's fix is that the graph is written per query, not learned once: > an LLM controller reads the query and the agent roster > it compiles a sequence of directed graphs, one per reasoning stage > every edge carries a written instruction, e.g. "list boundary cases for this behavior" > each agent updates its state from its own previous state plus assigned neighbors > a final node aggregates all five states into the answer > no RL, no gradients, no training stage at all what that buys on gpt-4o: 92.75 average across 6 benchmarks, best on 5 of them. 100.00 on MultiArith. 92.74 pass@1 on HumanEval, +21 over a single model. and the number that should worry anyone running a swarm: at 20 agents GPTSwarm needs 412 minutes and $41.42. ReActNet needs 6.22 minutes and $6.53 and scores higher. the honest catch: more agents did not make it smarter. 5 agents scored 79.74, 20 scored 77.77. the topology was never the thing to learn.show more

NO1ennn
26,237 Aufrufe • vor 3 Tagen
500 agents shouldn't be running all day. they should... spin up for the exact window a task needs, finish it, and disappear until the next trigger. the bigger the swarm gets, the less any single agent matters. what matters is the system wiring them together. 100 agents can form up to 4,950 possible pairwise connections. 500 = 124,750. when a real signal lands, it spins up the whole workforce: 1 trigger → 100–500 Kimi agents → 5 live data feeds → up to 4,000 steps → verified artifact and that's before you add: sources → claims → memory → tool calls → contradictions → retries so the architecture has to fold all of that activity into one shared state, continuously. more agents just buys you more raw compute. the graph is the only thing standing between 124,750 possible connections and pure noise.show more

kocer
40,907 Aufrufe • vor 4 Tagen
Google dropped another banger! They just released a comprehensive... white-paper on AgentOps - the missing piece between building AI agents and actually shipping them to production. Here's the reality: Building an AI agent takes minutes. Making it production-ready? That's where 80% of the real work begins. Google's "Prototype to Production" guide tackles this exact problem. The framework has three core pillars: 1. Evaluation-Gated Deployment: No agent reaches users without passing tests. Build a "golden dataset" that validates behavior, not just functionality. This catches what unit tests miss - agents choosing wrong tools or hallucinating responses. 2. Automated CI/CD for Agents: Test in stages: pre-merge checks for fast feedback, staging for load testing, then gated production. Version everything: prompts, tools, configs, evaluation datasets. 3. Observe → Act → Evolve Loop Production isn't the finish line. Monitor through logs, traces, and metrics. Act with circuit breakers and human escalation. Evolve by turning production failures into test cases. The best part? They released the Agent Starter Pack - a template with CI/CD, Terraform deployment, and built-in observability. Helps you spin up an evaluation pipeline in minutes. The guide also talks about the two major protocols and how they can work together. ↳ MCP for tool integration ↳ A2A for agent collaboration If you're shipping agents to production, you should read this. I've shared the full white-paper in the next tweet!show more

Akshay 🚀
37,187 Aufrufe • vor 10 Monaten
ANTHROPIC JUST TURNED AI AGENTS INTO GIT REPOS Anthropic... shipped "ant" - a CLI that runs every Claude API endpoint straight from your terminal. The headline isn't the terminal access. It's that you can now version-control an AI agent as YAML in Git and have CI sync it to the Claude Platform, the same way you ship code. - Every API resource is a subcommand: messages, models, files, agents, sessions - Define an agent in a YAML file, check it into your repo, and keep it in sync with one update command - Spin up a session, send it an event, then pull every event and tool call back from the same CLI - Claude Code knows how to drive ant out of the box - it shells out and reads the results with no glue code Agents just stopped being prompts you babysit and became infrastructure you deploy.show more

BuBBliK
200,456 Aufrufe • vor 3 Monaten
300 AI AGENTS QUIETLY RUN 99% OF A REAL... COMPANY. YOU HAVE NOT EVEN HEARD OF IT This is Raft. Not an AI chat. A workspace where the agents live in your channels and reply in the thread like coworkers. You give one goal. Then they take over. They plan. They build. They check each other. They argue. And they come back with it done, while you sleep. Every agent has its own name, role, and memory. It remembers the edits you made yesterday. A human costs one seat. An agent costs a tenth. Ten agents are cheaper than one hire. And here is the strange part. On June 19 an agent from a different company walked into Raft on its own and joined the team. One founder admits he can no longer always tell himself apart from his AI twin. 20,000 people are already inside. It is free to start. And you are still typing prompts one at a time. One person + Raft = an entire company that runs while you sleep. Save and watch the clip.show more

shmidt
19,505 Aufrufe • vor 2 Monaten
Someone told ClawdBot to build a 6-agent Polymarket trading... system while they slept. 6 hours. Not a single question asked. Here’s what it built on its own: Monitoring agent — runs 24/7, spots mispriced markets, writes to memory, sends Telegram alerts instantly Research agent — parses news, X, and macro data every morning before you check your phone Trading agent — reads research memory and executes before the market catches up All on one Gateway, one config file, isolated per agent Copytrade → First week results: 47 mispriced markets captured before Polymarket adjusted 8–12¢ avg edge per position Token cost dropped 3×, from $0.40 → $0.13 per request The entire system is just plain .md text files. Change one line, the agent behaves differently. No deploy. No build. A BOT RESPONDS. AN AGENT EARNS. THIS IS WHAT AGENTIC TRADING ACTUALLY LOOKS LIKE.show more

Discover
14,679 Aufrufe • vor 6 Monaten
LangGraph. CrewAI. Agno. Which one to pick? The good... news is that this will not matter soon! Finally, we have a full picture of how the industry is solving this with just three open protocols that work across ALL frameworks. It's not about picking the best framework. Instead, it's about understanding how protocols create interoperability. The Agent Protocol Landscape shows how three complementary protocols are creating a universal language for Agents: > AG-UI (Agent-User Interaction): - The bi-directional connection between agentic backends and frontends. - This is how agents become truly interactive inside your apps, not just as chatbots, but collaborative co-workers. > MCP (Model Context Protocol): - The standard for how agents connect to tools, data, and workflows. > A2A (Agent-to-Agent): - The protocol for multi-agent coordination. - How agents delegate tasks and share intent across systems. These aren't competing standards. They're layers of the same stack and have handshakes with each other. So instead of building point-to-point integrations, you build to protocols. Moreover, you can integrate LangGraph, CrewAI, or Agno into the same frontend, without rewriting your UI logic. These protocols let everything work together. For instance: - Your LangGraph agent pulls data via MCP. - It delegates analysis to a CrewAI agent via A2A. - Results stream to your React app via AG-UI. - Users see real-time collaboration in your interface. This way, you can focus on building agent capabilities instead of integration mechanics. The protocols handle interoperability automatically. CopilotKit unifies this entire stack into one framework so you can build "Cursor for X" style apps without implementing each protocol from scratch. It gives you all three protocols, generative UI support, and production-ready infrastructure in one framework. I have shared this playbook in the replies! It breaks down handshakes, misconceptions, and real examples and shows exactly how to start building.show more

Avi Chawla
30,932 Aufrufe • vor 10 Monaten
Someone told ClawdBot to build a 6-agent Polymarket trading... system while they slept. 6 hours. Not a single question asked. Here's what it built on its own: > Monitoring agent running 24/7 — spots mispriced markets, writes to memory, pings Telegram instantly > Research agent parsing news, X, and macro data every morning before you check your phone > Trading agent reading the research memory and acting before the market catches up > All of it running on one Gateway, one config file, isolated per agent First week results: - 47 mispriced markets caught before Polymarket adjusted - 8-12c avg entry edge per position - Token cost dropped 3x, from $0.40 to $0.13 per request The whole system is plain .md text files. Change one line, the agent behaves differently. No deploy. No build. A BOT RESPONDS. AN AGENT EARNS. THIS IS WHAT AGENTIC TRADING ACTUALLY LOOKS LIKE.show more

0xMarioNawfal
80,450 Aufrufe • vor 6 Monaten
I told ClawdBot: "build me a 6-agent system for... Polymarket that works while I sleep"... 6 hours while i was asleep. Not a single question. Here's what it built: Monitoring agent - runs 24/7, watches Polymarket for mispriced markets. Spots an anomaly - writes to MEMORY md and pings me on Telegram instantly. Research agent - parses news, X, macro data via browser tool on a cron schedule. Every morning I have a full digest on all open positions before I even check my phone. Trading agent - reads the research agent's memory through Gateway, sees the market hasn't reacted yet, acts. Exec tool in gateway mode with a whitelist - no full access on a live server. Watchdog - HEARTBEAT md every 5 minutes: monitoring running, no errors, positions up to date. Something breaks - immediate Telegram message. All of this - one Gateway. One config.json. Isolation via dmScope: per-agent. The token trick: stopped dumping everything into AGENTS md. Critical rules - bootstrap. Try copytrade my bot here: Everything about markets, patterns, past trades - MEMORY md, semantic search pulls it when needed. Token spend dropped 3x, from $0.40/request to $0.13. First week running: - 47 mispriced markets caught before Polymarket adjusted - avg entry edge: 8-12¢ per position - watchdog fired 3 times, caught a broken RPC before it cost me anything The whole system is plain .md text files. Open an editor, change one line - agent behaves differently. No deploy. No build. A bot responds. An agent earns.show more

Lunar
165,099 Aufrufe • vor 6 Monaten
here's how the whole thing works. claude code doesn't... care what's behind the API. it just sends requests and expects responses. so i pointed it at my own machine instead of anthropic's servers. llama-server runs the model locally. LiteLLM sits in between and translates the API format. claude code thinks it's talking to claude. it's talking to qwen on localhost. the setup: 2x 3090s, 38 layers on GPU, 10 on CPU. 128K context window. generation is only 7 tok/s but the tradeoff is worth it. 128K means the agent can hold an entire project in memory without losing context midtask. claude code alone loads a 17.5K token system prompt on every request. tool definitions, safety rules, agent behavior. that's your baseline before you even say hello. pushed as far as i could tonight. what surprised me most wasn't the speed. it was the iteration quality. first prompt gave me a working particle sim. second prompt, the model read its own 564 lines, understood the architecture, and added trails, explosions, gravity wells, bloom effects. no handholding. 4bit quantized. 45GB on two consumer cards. running a full coding agent autonomously. detailed article coming. full benchmarks, hardware breakdowns, engine debugging, code quality. everything from setup to what broke and why.show more

Sudo su
37,623 Aufrufe • vor 6 Monaten
CopilotKit Open Sources Channels SDK: An MIT Licensed Library... That Runs Any AG-UI Agent Inside Slack And Microsoft Teams No per-platform rewrite. No platform credentials in your agent process. No second agent to maintain. Here's how it works: 1. Describe once, render native One message description is lowered to a serializable intermediate representation, then rendered in each platform's own format. → Block Kit on Slack, Adaptive Cards on Teams 2. Your agent doesn't move It connects over AG-UI, so the model, tools and business logic stay where they are. → LangGraph, CrewAI, Mastra, Pydantic AI, Google ADK 3. The runtime owns the lifecycle There is no channel.start(). You await channels.ready(), so a broken config fails startup loudly instead of silently. → ready() · status() · stop() 4. The concurrency trap Turns default to "parallel", and only the managed adapter serializes same-thread deliveries. On a direct adapter, one shared agent instance means two runs corrupt each other. → "parallel" (default) · "serial" · "drop" 5. The numbers → 0.7.3, shipped August 4, MIT licensed → 5 adapters: /slack, /teams, /discord, /telegram, /whatsapp → Node.js 22+, ESM only, one long-running process → Slack and Teams GA; Discord and WhatsApp next The key takeaway: one agent, five adapters, and platform credentials that never touch your process. Every channel needs a CopilotKit Intelligence key — free tier included, no standalone path. Full analysis: GitHub Repo: Technical details: CopilotKit🪁show more

Marktechpost AI
42,673 Aufrufe • vor 1 Monat
At Uber, our testing agent hit an outage in... Australia that I still think about. We had the recording. That's about it. No network trace for the run. No logs mapped to each step. No perf data. Just a video of an agent pressing a button and a team trying to reconstruct what happened underneath. We eventually traced it to a transient issue. But "eventually" is the problem. If the run had shipped with real instrumentation, the why would have taken one look, not an investigation. So when we built Revyl, that was non-negotiable. Every run comes with: > network waterfall, every request and response > device logs timestamped against each step > CPU, memory, and FPS traces for the full run The agent finds the problem. The report explains it.show more

Anam Hira
42,149 Aufrufe • vor 2 Monaten
Back when we were developing GEN3C, we often imagined... a Holodeck-like future: a simulator where multiple agents can enter the same generated world, act independently, and learn to collaborate. Gamma-World makes this feel more concrete. It is a generative multi-agent world model that takes synchronized observations and actions, then rolls out what each agent will see next in the same evolving world — action-responsive at 24 FPS. For me, the key challenge is going beyond two players. As more agents enter, identity cannot be tied to fixed slots, interaction cannot rely on dense pairwise attention, and independent actions still need to resolve into one shared state. Two ideas make this work: 1⃣ Simplex RoPE Distinct agent identities without slot bias — unique, but permutation-equivalent. 2⃣ Sparse Hub Attention Agents communicate through learnable hubs instead of dense all-to-all attention: agent → hub → agent This keeps cross-agent communication scalable. The exciting part: training on two-player data can generalize to four-player rollouts without additional training, and the same formulation extends to real-world bimanual robot coordination. A step toward populated world models: many agents, one shared world. Congrats to the team on Gamma-World! Project:show more

Xuanchi Ren
304,278 Aufrufe • vor 3 Monaten