Загрузка видео...
Не удалось загрузить видео
my number 1 agent tip is "make the agent prove it understands things". everything boils down to that. the UVs on this book are the best example I have. if you don't do this and then ask the model to do something like "put some text on the front... show more
30,925 просмотров • 4 месяцев назад •via X (Twitter)
Комментарии: 16

you don't want to have to add this kind of debugging later, after things are going wrong. you want to build everything, from the beginning, so that it has some way to rapidly bridge agent understanding with human confirmation of this understanding.

now "prove it understands in a way I can verify" gets more and more abstract as you move out of stuff like 3D models and UVs and your own ability to come up with good ways of doing this will be the limiting factor, but the more you think like this the better you'll get at it.

can you say more about this? do you have the agent write tests? do you have some kind of tooling that you have built out to help you?

tests are part of it, but this is more like how you make sure the tests are testing the thing the agent thinks it's testing. for now the tooling is more lazy/custom per-project. partially because I'm building fake things with no stakes and just kind of figuring out what works as I go. this is a decent example: now, these are still kind of "fake". I didn't actually build this carefully and verify every test. but I sort of laid out the plan for how to build these tests with the visual verification in the way that I would want to IF I was building something that mattered, but then kinda skipped the due-diligence on my part aspect because I was just trying to build it as fast as I could with a reasonable amount of trust from the tests. but I think it gestures in the right direction for what a custom "dashboard" for this kind of verification looks like.

"this is more like how you make sure the tests are testing the thing the agent thinks it's testing" can you say more about this perspective? do you sort of have a sense of what kinds of things can go wrong? and, how much testing does the agent do on its own / do you do?

yes, you get used to the kinds of things that can go wrong. it is fairly intuition-based. but things that depend on a chain of different things that all how to work right are especially fragile. like in this case, text on a mesh in a specific place. it has to deal with so many different coordinate systems and keep track of converting between them. if you do this naively you'll get all kinds of problems. text that's upside down or backwards, then you ask it to fix that and you'll get text that's oriented right but now the top of the text is aligned instead of the bottom, or something like that. when that starts happening, the model is just throwing stuff at the wall, it doesn't *understand*. it only understands if it builds in that understanding from the beginning. it gets sort of crystallized when you build it right from the ground up, one step at a time, in variable names and tests and it gets harder for it to drift from it. the agent does build the tests, you just have to make sure the errors aren't cancelling out in the test. like say you ask it to confirm text layout with a test. how do you know this test path matches the path in the actual code? what if the output is flipped and this makes the text oriented correctly, but in the live scene the text is upside down? you have to make sure that you start with a foundation that doesn't have these kind of errors so you can trust the tests it builds on top of them. with this book example when I say "write a test that makes sure left/center/right alignment for text works", I'd trust that it's testing the right thing, because it knows what up/down/left/right are in the space of the surface of the front cover of the book. even if there was an issue with the text sampling so each glyph somehow ended up horizontally flipped, and I said "the text is horizontally flipped", I'd expect it to actually fix the right thing, instead of potentially trying changing some other coordinate space and creating more problems.

Making tasks verifiable and using verification as part of ongoing design is *really useful* beyond and even without rl and so on

Yep. Prove grounding first then let it act. otherwise you are just asking latent mush to improvise geometry

very nice

@charliermarsh uv mentioned

Good example.

was that 3d model created by an ai model? Which one?

yeah same, while working in unity I always ask for variety of gizmos previews

@soulblocks Have an example you used ?

Honestly, I treat UVs as structure. If the agent can't map the seam between cover and spine, it's not modeling. It's guessing with confidence.

what's this video from

