正在加载视频...
视频加载失败
My Ralph setup has evolved a LOT since you last saw it. It's totally AFK, closing GitHub issues while I work on my courses. I've learned a lot along the way. Here's a breakdown:
54 条评论

I'm running a live workshop! 11th Feb 9AM-1PM PST 40 attendees max Learn how to get the most out of Ralph

Sold out! Thanks folks

Interesting. My setup is quite different and I emphasized on more smoke tests using playwright and/or end to end tests + screenshots. Do you always read the activity log file in an iteration? I've been doing that but as you noticed it's starting to get fairly large. I am thinking maybe it can be optional "if uncertain about functionality review the implementation log" so it doesn't always put it into context. That's currently the elephant in the room since it clearly won't scale infinitely and make it more dumb as we implement tasks 🤔

What I should be doing with Progress.txt is saving it to some external database and taking only the last 10 entries. What we want it for is not necessarily a log for the project but a log for the type of tasks that the agent has performed to give it some context as to the type of things that it's done and the size of tasks it's attempted. In other words, it's a kind of smoothing function.

@d4m1n For this, I have Ralph create a new progress. md and notes md - per prd, so it doesn't stack in one. I then delete those after each task list is complete.

@d4m1n If you're curious, I use this for Ralphing

@mattpocockuk thanks for sharing! Very interesting 🙏

I see Pocock post, I bookmark then watch

ahh back a week ago or so when i complained it was boring waiting at a blank terminal you sent me to touch grass lol.. good stuff!

😂

Can you link us to the repo?

Thank you!

AFK that actually closes issues is the line between “toy agent” and “teammate”. The big unlock is you’re turning work into a repeatable loop: tight inputs, visible outputs, and a clear definition of done.

No one is adopting that stupid AFK name no matter how hard you keep trying to make people

What are you talking about

The goat

Would you say this approach is most useful for adding new features / tweaking existing projects? And detailed PRDs still needed for spinning up new stuff with Ralph? Curious to know if / when you still use the PRD.json approach

Nope I don't use the .json right now, but not averse to it .md is more flexible and easier to read for humans PRD's go in the backlog, Ralph eats them up same as issues

this works really well 👍

Most value comes from refining workflows over time.

I’ve been having great luck with amp/claude pulling from @linear, moving to in progress, working until complete, and grabbing another. All while sleeping 😴 WILD times!

Crisp audio what mic are you using? Also thanks for the update!

MV7!

Is it possible to get the browser feedback loop working when using Ralph in docker sandbox?

Yes, chrome devtools MCP with --no-sandbox works great for me

Do you have chrome installed in the docker sandbox?

Hey one of your last videos inspired me to hack this together over the last couple of days - it's an orchestration framework that runs work items in parallel - works with github issues or plain plans:

Maybe if progress.txt is written in XML format Claude Code can actually save context every time it needs to get information from there, I think.

Curious about the guardrails : any 'oh shit' moments with it running unsupervised yet?

Interesting topic, thnx for sharing it

how does it know it actually completely work if it's just running tests and typechecking? are you putting full PRDs and test steps into the github issues?

any ralph on google gemini cli or anti-gravity lol , i stop using claude code for a very long time

agents should at the very least burndown your tickets for you 24/7 but probably in serial parallel execution is where you get into trouble

When you QA everything, are you reviewing code? Or just confirming functionality? I love the concepts behind Ralph but I find it’s too easy to generate a ton of code. Straight into vibe coded app territory.

A bit of both. Code quality is ESSENTIAL for keeping Ralph healthy.

The Ralph evolution is insane. AFK automation while you build is the dream setup 🔥

auto triage plus templated fixes is seriously clever

Excellent

@grok wha is Ralph?

do you use claude to generate the issues with the anthropic recommended format to begin with?

have you tried using beads? is that not better?

Thanks for all the Ralph lately! Question: Say you find a bug in what the agent implemented for a bigger PRD. How do you get them to track and fix it? Do you add another item to the PRD, or bail back into a less Ralph-y coding workflow?

this helped me a ton with that json nonsense you were dealing wiht

Is there any significant difference between Tracer bullets and vertical slices?

I like it because it's a well-known programming axiom, and helps tickle the right latent space for AI to behave properly

But in practical terms, no

Pretty genius idea to just connect the issues via the gh cli. I wonder if making it run in parallel with worktrees would be a good idea. Furthermore I could see issue tag being useful to "ready" tasks to work on for ralph.

Yeah maybe, I'm imagining that running multiple Ralphs will basically be a queue management exercise

Was just running this setup. It's amazingly simple and effect. Was telling it to work on a branch and create a PR when finished. Felt a bit off coming back to have 10+ commits to main. But reviewing PRs can be too cumbersome. What is appropriate depends on the project and team.

@mattpocockuk You can use anthropic's or openai GitHub integrations to run code reviews. It adds a GitHub workflow and after the review they comment automatically to the PR.

@mattpocockuk So I have todo a Review of the Review? Not sure if Reviews are helpful at all anymore. Heavily depends on the project for sure. This fast AI coding is more fire commits and tight feedback loops. Reverting on breaks. Extensive TDD and other checks is the way for me imho.

Awesome - but the real win is the safety rails: auto-close only when it can reproduce + attach failing/passing test evidence; otherwise label/triage. What are your "never touch" rules (security/infra/payments)?

Banger as always Matt! 🙏🏽👌🏼

Have you found ralph to be useful still after CC released tasks?

