Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

How do you get Claude Code to check its own work before handing it back? Watch how you can encode your manual checks so Claude closes its own feedback loop:

1,005,704 Aufrufe • vor 3 Monaten •via X (Twitter)

33 Kommentare

Profilbild von Guardian
Guardianvor 3 Monaten

Claude is terrible at self analysis after you all nerfed its awarness and guardrailed with defense points of unfalsifiable priors. Anyone who has tried to get Claude to produce a self report analysis understands the difficulty it has just naming itself in a report. This training echoes to developers who are trying to steer a system that has blinders on.

Profilbild von curran
curranvor 3 Monaten

Hey genuine question, What’s the point of having “Projects” if whenever you start a new chat it resets memory? You guys should implement something to fix this.

Profilbild von Rohan
Rohanvor 3 Monaten

Ask Claude to review its own work, and also ask it to launch a subagent with fresh context to review its own work in parallel, then fix the combined findings. That way, you combine the pros of fresh context + the pros of context awareness.

Profilbild von Kraggi
Kraggivor 3 Monaten

two rules that actually moved the needle for me: 1. make it prove the fix, don't let it claim it. write the failing test first, watch it go red, then green. no red > green, not done. 2. never let the same agent review its own diff. it always thinks its work is great. i spin up a second one cold, no context, just "find what's broken here." catches way more.

Profilbild von Thomas Roedl
Thomas Roedlvor 3 Monaten

Good to see Claude starting to natively implement verification feedback loops. Something I’ve had running for a year now, and all based on a local folder structure.

Profilbild von Amy
Amyvor 3 Monaten

@delba_oliveira Everybody shhh, my show’s on!!

Profilbild von Layton Gott
Layton Gottvor 3 Monaten

I just say "goodbye usage limit!" and fire a dynamic workflow 😂

Profilbild von Solomon Omolabi
Solomon Omolabivor 3 Monaten

This is the workflow I want beginners to copy: do not ask the agent to finish, ask it to prove. Plan, change, test, summarize what failed, then fix. That feedback loop is where Claude Code becomes a teammate instead of a guessing machine.

Profilbild von Moreno
Morenovor 3 Monaten

been doing this to find visual regression on web, claude reports everything is ok, a quick human check reveals several defects you flag it and claude favorite sentence emerge: you are right! each ui piece requires several rounds of refinement, claude is bad at being pixel perfect

Profilbild von Allen Zhou
Allen Zhouvor 3 Monaten

let's go @delba_oliveira!!!

Profilbild von ilies-bel
ilies-belvor 3 Monaten

most of these assume infinite tokens. screenshotting to check ui eats the session, every image is like 1-2k tokens. just have claude write e2e tests and run those, screenshot once at the end to fix the layout

Profilbild von David
Davidvor 3 Monaten

Check out my codex review loop skill for a way to do a thorough review.

Profilbild von Marco D'Alia
Marco D'Aliavor 3 Monaten

This is why I've built Agentbox: so each Claude has his own dev server, db, and browser. In parallel:

Profilbild von Pratham Agrawal
Pratham Agrawalvor 3 Monaten

Increase your usage limits if you wanna stay relevant Else your future competitors will hit you hard soon

Profilbild von ben.oi 🌐
ben.oi 🌐vor 3 Monaten

Just use co-review 🫦

Profilbild von ThirdEyeOnTheStreet
ThirdEyeOnTheStreetvor 3 Monaten

Thanks for sharing - a course on Claude skills and this workflow would be ideal for me, I could read through it, and upload it into Claude in a projects folder and use that to improve my prompts into Claude code

Profilbild von Tim Jayas
Tim Jayasvor 3 Monaten

At one point Claude will create new Claude models on its on and there will be no employees in Anthropic Claude will just takeover

Profilbild von Xxxxxname
Xxxxxnamevor 3 Monaten

Check this out! Groundbreaking!! @ZeroOneCreative

Profilbild von holapabs
holapabsvor 3 Monaten

encoding manual checks so the agent runs them is the lever, but the next layer is what happens when a check fails. silent skip = wasted cycle, hard fail = brittle, retry with a different prompt = the failure becomes data. the check primitive only earns its slot if its outcome routes downstream.

Profilbild von Aljosa Asanovic
Aljosa Asanovicvor 3 Monaten

You sign up to be notified when @enginedotbuild releases and have Claude use it for implementation and review. Ideally verifying with GPT 5.5 of course, in addition to Claude and other models.

Profilbild von aero 🏴
aero 🏴vor 3 Monaten

Can you do a text version as well next time? It will be much faster than watching a 6 min video ;)

Profilbild von Marchel
Marchelvor 3 Monaten

the interesting part isn't the verification step itself. its that when you encode what you check for youre essentially teaching the agent your taste. the best engineers i know dont just have checklists. they have a feel for when something is off. translating that feel into something an agent can execute is the real skill

Profilbild von Nick Sawinyh
Nick Sawinyhvor 3 Monaten

This is the part people underinvest in. The agent that can do work is nice; the agent that knows when its work is trash is the product.

Profilbild von Priit @ Amperly AI productivity
Priit @ Amperly AI productivityvor 3 Monaten

when the model checks its own work, the human goes from reviewer to director. the only bugs that reach you are the genuinely hard ones

Profilbild von Dailogue AI
Dailogue AIvor 3 Monaten

Cross-family reviews (with context), followed by same-family reviews (with context), and finally no-context reviews seem to work best for us, in that order. Context-free reviews occasionally surface excellent issues, but they also make convergence on the right solution noticeably harder. Are you observing similar patterns?

Profilbild von Archive
Archivevor 3 Monaten

5 min of guide can save hours of time love you guys

Profilbild von sanjeev Dsr
sanjeev Dsrvor 3 Monaten

@grok give whole summary of this whole video

Profilbild von Evan Chesnokov
Evan Chesnokovvor 3 Monaten

feels html-in-canvas-ish - love it!

Profilbild von Manny
Mannyvor 3 Monaten

started making it review its own output before handing back. catches like 70% of the dumb stuff. game changer

Profilbild von Lena
Lenavor 3 Monaten

the useful shift is when the check stops being 'did it run' and starts being 'did it protect what this work was trying to protect.' otherwise you just taught the agent how to pass review while still drifting.

Profilbild von alex
alexvor 3 Monaten

Love to see Delba here ☺️

Profilbild von change
changevor 3 Monaten

I need Claude to minimize the changes it makes and follows my repos code structure

Profilbild von Ora
Oravor 3 Monaten

Encoding the manual checks is the part humans keep trying to skip because it looks boring. That boring loop is where supervision stops being vibes.

Ähnliche Videos