ๆญฃๅœจๅŠ ่ฝฝ่ง†้ข‘...

่ง†้ข‘ๅŠ ่ฝฝๅคฑ่ดฅ

๐—œ ๐—ผ๐—ฝ๐—ฒ๐—ป ๐˜€๐—ผ๐˜‚๐—ฟ๐—ฐ๐—ฒ๐—ฑ ๐—บ๐˜† ๐—น๐—ผ๐—ฐ๐—ฎ๐—น ๐—”๐—œ ๐—ฝ๐—ฒ๐—ฟ๐˜€๐—ผ๐—ป๐—ฎ๐—น ๐—ฎ๐˜€๐˜€๐—ถ๐˜€๐˜๐—ฎ๐—ป๐˜: ๐—”๐—ด๐—ฒ๐—ป๐˜ ๐—›๐—ฎ๐—ฟ๐—ป๐—ฒ๐˜€๐˜€ & ๐—Ÿ๐—ผ๐—ผ๐—ฝ ๐—˜๐—ป๐—ด๐—ถ๐—ป๐—ฒ๐—ฒ๐—ฟ๐—ถ๐—ป๐—ด ๐—ถ๐—ป ๐—ฟ๐—ฒ๐—ฎ๐—น ๐—ฐ๐—ผ๐—ฑ๐—ฒ. Everything in real code you can read in an afternoon: Harness, Loop Engineering, Memory, Eval, Tracing. 100% local on your laptop. ๐Ÿ’ป Github Repo: โ˜•๏ธ Buy me a coffee: The whole system is...

46,123 ๆฌก่ง‚็œ‹ โ€ข 2 ไธชๆœˆๅ‰ โ€ขvia X (Twitter)

34 ๆก่ฏ„่ฎบ

Shen Sean Chen ็š„ๅคดๅƒ
Shen Sean Chen2 ไธชๆœˆๅ‰

Full Tutorial:

Shen Sean Chen ็š„ๅคดๅƒ
Shen Sean Chen2 ไธชๆœˆๅ‰

Github Repo: โ˜•๏ธ Buy me a coffee: Star it if it's useful. PRs welcome, the voice mode especially. ๐Ÿ™

morluto ็š„ๅคดๅƒ
morluto2 ไธชๆœˆๅ‰

i found this to be a really engaging way to show the agent loop!

Shen Sean Chen ็š„ๅคดๅƒ
Shen Sean Chen2 ไธชๆœˆๅ‰

Thank you. Try it yourself. Would love to hear any feedback.

Shen Sean Chen ็š„ๅคดๅƒ
Shen Sean Chen2 ไธชๆœˆๅ‰

I built an arena in Waku Agent to compete Kimi K3 with Claude Opus 4.8, Fable 5, Grok, GPT, Gemini. The results are fun. Clone the repo and try it out yourself:

Shen Sean Chen ็š„ๅคดๅƒ
Shen Sean Chen2 ไธชๆœˆๅ‰

๐—ช๐—ฎ๐—ธ๐˜‚-๐—”๐—ด๐—ฒ๐—ป๐˜ ๐—ถ๐˜€ ๐—ป๐—ผ๐˜„ ๐—ถ๐—ป๐˜๐—ฒ๐—ด๐—ฟ๐—ฎ๐˜๐—ฒ๐—ฑ ๐˜„๐—ถ๐˜๐—ต @OpenRouter, ๐˜„๐—ต๐—ถ๐—ฐ๐—ต ๐—ฟ๐˜‚๐—ป๐˜€ ๐—ฏ๐—ฎ๐˜€๐—ถ๐—ฐ๐—ฎ๐—น๐—น๐˜† ๐—ฎ๐—ป๐˜† ๐—บ๐—ผ๐—ฑ๐—ฒ๐—น ๐˜†๐—ผ๐˜‚ ๐˜„๐—ฎ๐—ป๐˜. ๐—ข๐—ป๐—ฒ ๐—ธ๐—ฒ๐˜†, ๐—ต๐˜‚๐—ป๐—ฑ๐—ฟ๐—ฒ๐—ฑ๐˜€ ๐—ผ๐—ณ ๐—บ๐—ผ๐—ฑ๐—ฒ๐—น๐˜€. (220 stars in 2 days: Shipped as a community PR from maanas1234, tested against the eval gate, merged. Thank you! Try it: WAKU_PROVIDER=openrouter WAKU_MODEL=meta-llama/llama-3.1-70b-instruct (or any model slug) You Can Build Anything. You Can Learn Anything. ๐Ÿ’ช

Hades ็š„ๅคดๅƒ
Hades2 ไธชๆœˆๅ‰

Avg 10+ commits per day. Man, I know you love it so much, respect ๐Ÿซก

Shen Sean Chen ็š„ๅคดๅƒ
Shen Sean Chen2 ไธชๆœˆๅ‰

๐Ÿซก๐Ÿซก

Rohit ็š„ๅคดๅƒ
Rohit2 ไธชๆœˆๅ‰

That's awesome. Glad to see you pushing the boundaries with open source. Nicely done.

Shen Sean Chen ็š„ๅคดๅƒ
Shen Sean Chen2 ไธชๆœˆๅ‰

@ai_rohitt Thanks

kai Nakamura ็š„ๅคดๅƒ
kai Nakamura2 ไธชๆœˆๅ‰

Tracing is the piece most harness diagrams leave off, and it's the one that makes the rest usable. Without a record of what the loop did and why, memory and eval are just guessing. Readable-in-an-afternoon is the right bar too: if you can't read the harness, you can't trust it.

Scott Bonge ็š„ๅคดๅƒ
Scott Bonge1 ไธชๆœˆๅ‰

Got a video like this but geared toward increasing sales by images generating and posting new content instead of customer service? Iโ€™m a small business owner โ€” invented my product, been selling 20+ years online. Customer service isnโ€™t my bottleneck; getting more eyes on the product and driving sales is. Curious if youโ€™ve done one on setting up an agent for that (e.g., generating product content and posting to social) rather than the customer service angle?

AI Mastery Guide ็š„ๅคดๅƒ
AI Mastery Guide2 ไธชๆœˆๅ‰

95 lines for a full agent loop is genuinely impressive

Shen Sean Chen ็š„ๅคดๅƒ
Shen Sean Chen2 ไธชๆœˆๅ‰

cheers

AI Mastery Guide ็š„ๅคดๅƒ
AI Mastery Guide2 ไธชๆœˆๅ‰

๐Ÿ‘

murphy ็š„ๅคดๅƒ
murphy2 ไธชๆœˆๅ‰

Thats amazing๏ผthank you for sharing

Shen Sean Chen ็š„ๅคดๅƒ
Shen Sean Chen2 ไธชๆœˆๅ‰

Youโ€™re welcome! Would love to hear any feedback if you tried it.

C.G. ็š„ๅคดๅƒ
C.G.1 ไธชๆœˆๅ‰

You get a donation! lol please keep making these, this one and the agentic loop vid were ๐ŸชŽ

Shen Sean Chen ็š„ๅคดๅƒ
Shen Sean Chen1 ไธชๆœˆๅ‰

Thanks so much ๐Ÿซถ

Ankush Singal ็š„ๅคดๅƒ
Ankush Singal2 ไธชๆœˆๅ‰

thanks for sharing , are you directly using tools or you created MCP server?

Shen Sean Chen ็š„ๅคดๅƒ
Shen Sean Chen2 ไธชๆœˆๅ‰

both, in the current repo itโ€™s mainly tools. But more MCP connectors coming soon.

Ankush Singal ็š„ๅคดๅƒ
Ankush Singal2 ไธชๆœˆๅ‰

thank you, great work!!

Shen Sean Chen ็š„ๅคดๅƒ
Shen Sean Chen2 ไธชๆœˆๅ‰

Thanks

qtieepiee ็š„ๅคดๅƒ
qtieepiee2 ไธชๆœˆๅ‰

This is the right local-agent stack: loop, harness, tracing, evals, memory, and skills. SQLite memory is a great start, but lifecycle is where it compounds. Atomic Memory is worth checking out for that layer

Shen Sean Chen ็š„ๅคดๅƒ
Shen Sean Chen2 ไธชๆœˆๅ‰

did u build his?

Qin Chi ็š„ๅคดๅƒ
Qin Chi2 ไธชๆœˆๅ‰

Great open source. The harness + loop engineering approach for a local PA is super practical โ€” would love to dig into the memory/eval parts.

MPT Pay ็š„ๅคดๅƒ
MPT Pay1 ไธชๆœˆๅ‰

I can help with Voice, I have amazing quality for AI PSY HELP, pls try in the voice chat

dhinna ship .ico ็š„ๅคดๅƒ
dhinna ship .ico2 ไธชๆœˆๅ‰

Loop harnesses read clean until a bad tool call poisons memory. Tracing rarely catches the drift before evals pass

DX ็š„ๅคดๅƒ
DX1 ไธชๆœˆๅ‰

Such a clear demonstration and explanation. I am watching lots of your videos today. They really help me get a clear picture of Harness and Memory.

Stair AI ็š„ๅคดๅƒ
Stair AI2 ไธชๆœˆๅ‰

This is a great thing to open source. The 95-line loop and "open the sqlite file, it's yours" especially, most people never get to see the whole thing in one readable piece. thanks for putting it out there, starred.

Michael Barrick ็š„ๅคดๅƒ
Michael Barrick1 ไธชๆœˆๅ‰

does this only work via desktop vs VPS??

Jared Traehorn ็š„ๅคดๅƒ
Jared Traehorn1 ไธชๆœˆๅ‰

SQLite is the honest choice. the question is what happens when the loop writes state that contradicts 3 sessions ago and the older version was right. no rollback means trusting the last write forever.

Aryaman Upmanyu ็š„ๅคดๅƒ
Aryaman Upmanyu2 ไธชๆœˆๅ‰

having the entire agent loop open and readable in under one hundred lines of python is a dream for local builders

Eddie ็š„ๅคดๅƒ
Eddie2 ไธชๆœˆๅ‰

Separating deterministic tool checks from judged response quality creates a useful eval boundary. The next runtime question is escalation when a live turn becomes uncertain does the harness ask stop or keep acting That policy matters as much as the release gate.

็›ธๅ…ณ่ง†้ข‘

Hermes agent just left the terminal. ๐—›๐—ฒ๐—ฟ๐—บ๐—ฒ๐˜€ ๐——๐—ฒ๐˜€๐—ธ๐˜๐—ผ๐—ฝ dropped yesterday. native app for macOS, Windows, and Linux. for months Hermes was the agent that learned your projects, wrote its own skills, and built a model of who you are. all of it buried in terminal logs. now it has a window. the important part is that it's not a wrapper. it runs the same agent core, the same sessions, memory, and skills as the CLI. you can start a task in the terminal and finish it in the app without anything resetting. the state is shared across every interface, not copied between them. what the GUI actually adds: โ†’ streaming chat that shows live tool calls and inline reasoning instead of a spinner โ†’ a preview rail that renders pages, code, and images right beside the conversation โ†’ an artifacts panel that collects every file the agent has ever produced โ†’ remote gateway mode, so you can point the app at a VPS and run the heavy work elsewhere โ†’ skills, cron, profiles, and gateways managed point-and-click instead of through YAML โ†’ voice mode, drag-drop files, and inline image generation remote gateway mode is the one worth slowing down on. the agent runs 24/7 on a $5 server while you control it from your laptop like a local app. other agent UIs are chatboxes with a logo. this one shows the autonomy instead of hiding it, so you watch the skills load, the tools fire, and the artifacts pile up as it works. it was teased in Jensen's GTC keynote. MIT licensed, local-first, no telemetry. if you already run Hermes, download it and everything is already there. your chats, memory, and skills carry straight over. i wrote a full masterclass on Hermes Agent that walks through the SOUL. md identity layer, the three-tier memory system, the self-evolving skills loop, and how to run three specialized agents 24/7. desktop is the interface that finally does all of it justice. the article is quoted below.

Akshay ๐Ÿš€

51,540 ๆฌก่ง‚็œ‹ โ€ข 3 ไธชๆœˆๅ‰

HERMES AGENT LEARNS FROM ITS OWN MISTAKES. UPDATES ITS MEMORY. CREATES ITS OWN SKILLS. NO CLOUD. EVERYTHING STORED LOCALLY. THIS IS HOW THE SELF-IMPROVING LOOP WORKS. most agents start from zero every session. Hermes carries forward what it learned. THREE MEMORY SYSTEMS: 1. PROCEDURAL MEMORY (how to act) stored in ~/.hermes/skills/ as SKILL.md files. when the agent repeats a complex workflow, it saves the procedure as a reusable skill. next time the same task comes up, it follows the skill instead of figuring it out again. you can also create skills explicitly: "create a skill called video-prep that captures how I format my video scripts. spoken english, define jargon inline, no em-dashes, close with a catchphrase." the agent writes the SKILL.md. available as a slash command from that moment. Hermes ships with 90+ skills. the number grows the longer you use it. 2. SEMANTIC MEMORY (durable facts about you) stored in ~/.hermes/memory/memory.md the agent scans conversations for facts worth remembering. preferences, habits, corrections, project details. real example from the video: agent tried to scrape a YouTube channel. URL was wrong. it failed. it updated memory.md with the correct URL pattern so it never makes the same mistake again. you can also save explicitly: "save to memory that my favorite testing framework is pytest" the agent updates memory.md immediately. this file loads into context on every session. the agent knows you better every week. 3. EPISODIC MEMORY (chat history) stored in ~/.hermes/state.db (local SQLite). every conversation. every tool call. every result. searchable with FTS5 full-text search. "search our past sessions. what was the first thing I ever said to you?" the agent queries state.db and finds it. over time, auxiliary models consolidate episodic memory into semantic memory. distilling recurring patterns into durable facts. THE SELF-IMPROVING LOOP: every agent run follows this cycle: โ†’ you send a prompt โ†’ working memory loads: SOUL.md + memory.md + relevant skills + chat history โ†’ agent calls tools (terminal, browser, delegate_task) โ†’ agent completes the task, replies to you โ†’ AFTER the reply: agent checks "did I learn something worth saving?" โ†’ if yes: updates memory.md or creates a new skill โ†’ next session starts smarter than the last this happens automatically. you don't ask the agent to learn. it decides what to remember on its own. WHAT MAKES THIS DIFFERENT FROM CLAUDE CODE: Claude Code has memory too. but Hermes stores everything locally. no cloud. your data never leaves your machine. Claude Code doesn't auto-create skills from experience. Hermes turns repeated workflows into reusable procedures. Claude Code memory is instruction-based. Hermes memory is conversational and self-updating. over months of usage, Hermes builds a knowledge base of your preferences, your projects, your mistakes, and the procedures that work for your specific workflow. the agent that remembers your birthday also remembers why your last deploy failed. NO EMBEDDINGS. PLAIN TEXT. Hermes does not use embeddings or RAG for memory. skill and memory search runs on plain text keyword matching. simpler. faster. no vector database to maintain. works entirely offline on your local machine. DELEGATE TO CLAUDE CODE: Hermes can spawn a sub-agent that runs Claude Code in headless mode: "spawn a sub-agent using Claude CLI to build a Python script that fetches the top 5 Hacker News stories to markdown." Hermes delegates. Claude Code writes the code. result returns to Hermes. Hermes runs the script and delivers the output. use Hermes for orchestration. use Claude Code for heavy coding. both tools. not competitors. WHAT HERMES DOES NOT HAVE: no built-in eval or LMOps system. no LangSmith, no LangFuse integration out of the box. trajectory export and logs exist but there is no automated quality tracking. if you need eval, build it yourself or connect external tools. the loop is self-improving. measuring how well it improves is on you. comment LOOP and I'll send you the configs that control how fast Hermes learns and what it remembers. memory limits, skill auto-creation triggers, and the auxiliary model that runs the learning. Replace your entire team with 8 hermes agents๐Ÿ‘‡

YanXbt

22,720 ๆฌก่ง‚็œ‹ โ€ข 2 ไธชๆœˆๅ‰

Memory vs. Graphs, clearly explained! memory is great, and the ceiling arrives quietly: it stores what happened. it does not store what to do about it. six runs later your file has fifty lines, and the model reloads all of them before it does anything. Graph engineering fixes this by changing what memory is: not a place things are kept, but an edge that runs backwards. you need both, and here is the sentence that resolves the whole confusion: a store keeps what happened. an edge keeps what to do about it. โ†ณ a store grows with every run, and every line is reloaded before the next one โ†ณ an edge carries one derived rule, and the rule replaces the run that produced it Prompts โ†’ Context โ†’ Harness โ†’ Loops โ†’ Graphs the transcript goes away, the constraint stays. and the constraint is smaller, because "adapters preserve keyword args exactly" is four hundred tokens shorter than the run that proved it. the same four blocks work on anything you can cut into lanes. i pointed them at token launches on Robinhood Chain, open source, nothing leaves your terminal the trick is knowing what deserves to survive. an output is not memory. "ported the utils slice, green on first pass" tells the next run nothing it can act on. the rule you derived from it does. one thing to know before you scale it. what you write down is not what comes back. โ†ณ the root rules file and auto memory are re-injected from disk. they come back intact, every time โ†ณ path-scoped rules live in message history. they get summarized away and do not return until a matching file is read again so a rule that must persist cannot be path-scoped. move it to the root and pay the always-loaded cost, or accept that it is advisory in any long session. and the one that eats whole nights: a memory file that has never had a line deleted is not memory. it is a tax on every run you will ever make, and nobody reads it back. below i have quoted my full guide on graph engineering. it covers the three topologies, the verifier patterns, and where the gate should actually open. save this, and the repo that runs it is below โ†“

Hanako

47,766 ๆฌก่ง‚็œ‹ โ€ข 17 ๅคฉๅ‰

A finance professor manages $200M with AI agents, and he told everyone why: "Large language models are at the level of a fourth-year PhD student in every field" Alejandro Lopez-Lira's AI fund, Autopilot, returned 56% last year. The S&P did 16%. There are 52,000 people with money in it, and most of them just watch the machine work. What he automated is the same six-step loop every fund on earth runs: find an idea, code it, backtest it, deploy it, read the autopsy, learn from it. A quant at Two Sigma runs that loop once a month, and the salary time alone costs around $50,000 per hypothesis. All steps from this loop now fit in AI trading text box. Plain English in, executable strategy out, five-year backtest in 12 seconds, live on a broker 90 seconds after you typed the sentence. He runs $200M with AI. You can run same AI fund in two clicks, free to try: Step 6 on this loop is where everyone is stuck. Your agent has no memory. Every strategy it kills goes into a log nobody reads, and the next one starts from zero. Nobody keeps negative results. Not Citadel, not Man Group, not a single repo on GitHub. Fix that and the agent remembers every hypothesis it killed and the regime it died in. It stops burning cycles on your old mistakes. Jane Street pays 3,500 people to run this cycle and made $39.6 billion doing it. Five sixths of it is now free. Bookmark & read full map of this loop in the article below. Most people still think AI trading is out of reach for them - it isn't. Don't want to spend a dollar for testing this? Kalshi just opened a perps exchange and gives US users $25 free to start ->

cvxv666

83,045 ๆฌก่ง‚็œ‹ โ€ข 1 ไธชๆœˆๅ‰

A finance professor manages $200M with AI agents, and he told everyone why: "Large language models are at the level of a fourth-year PhD student in every field" Alejandro Lopez-Lira's AI fund, Autopilot, returned 56% last year. The S&P did 16%. There are 52,000 people with money in it, and most of them just watch the machine work. What he automated is the same six-step loop every fund on earth runs: find an idea, code it, backtest it, deploy it, read the autopsy, learn from it. A quant at Two Sigma runs that loop once a month, and the salary time alone costs around $50,000 per hypothesis. All steps from this loop now fit in AI trading text box. Plain English in, executable strategy out, five-year backtest in 12 seconds, live on a broker 90 seconds after you typed the sentence. He runs $200M with AI. You can run same AI fund in two clicks, free to try: Step 6 on this loop is where everyone is stuck. Your agent has no memory. Every strategy it kills goes into a log nobody reads, and the next one starts from zero. Nobody keeps negative results. Not Citadel, not Man Group, not a single repo on GitHub. Fix that and the agent remembers every hypothesis it killed and the regime it died in. It stops burning cycles on your old mistakes. Jane Street pays 3,500 people to run this cycle and made $39.6 billion doing it. Five sixths of it is now free. Bookmark & read full map of this loop in the article below. Most people still think AI trading is out of reach for them - it isn't.

sopersone

48,482 ๆฌก่ง‚็œ‹ โ€ข 15 ๅคฉๅ‰

Once you learn these three things, you can build nearly anything yourself. Skills, chains, and plugins. Learn them once and you can automate a real part of your week yourself. Here's the system: A skill is a standard operating procedure. It's one long, reusable prompt that does one thing well, like writing a newsletter, setting up a PPC campaign, triaging your inbox, or drafting a note to the board. You write it once and reuse it forever. The trick is to feed it your real work. For my writing skill, I gave Claude posts I admire and my own exported analytics with the winners marked, so it could see the patterns. Show it what good looks like and it nails your voice. Skip that step and it guesses. A chain connects skills into an automation. One skill's output feeds the next, in order. A copywriting skill writes the messaging, then hands it to a PPC-setup skill that builds the campaign. That's an automation, and you built it without writing code. A plugin is the harness that holds it all. It bundles your skills and chains into one thing you share with your team. Instead of sending 15 separate prompts around, you send one plugin, and it knows when to use each skill on its own. Skill, chain, plugin. It's one step up the ladder, and once you're up there it's easy. Now the worked example. I call it the Daily Driver. It's five skills chained together: email triage, a writer, Slack triage, a thinking partner, and a setup skill that connects my tools and personalizes everything. I chained them into one morning brief and scheduled it to run at 9am and message me the list. Two things made the biggest difference. Give it memory. I made plain docs for About Me, my Brand Voice, and my working preferences. I had Claude interview me and saved each answer as a file. Now every output sounds like me and knows my projects and my team. Connect your tools. The plugin reads my inbox, Slack, and calendar through MCP connectors for Gmail, Slack, and Google Calendar. That's what turns "summarize my morning" into a real brief instead of an empty wish. This is the highest-ROI thing an operator can set up this week. I run my YouTube channel, podcast, community, and newsletter on plugins I built myself, all on the $100 Max plan with no engineer in sight. I made a full walkthrough that shows the whole build start to finish. Want the Daily Driver plugin to start from? Comment PLUGIN below and I'll send it to you. #AI #automation #ClaudeCode

JJ Englert

53,969 ๆฌก่ง‚็œ‹ โ€ข 3 ไธชๆœˆๅ‰