Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Morning Bathrobe Rant: Rethinking Harnesses.

1,878,865 Aufrufe • vor 4 Tagen •via X (Twitter)

35 Kommentare

Profilbild von Behzad
Behzadvor 4 Tagen

alternative title: uncle bob vs the bitter lesson

Profilbild von LogosCat
LogosCatvor 4 Tagen

I told it several months ago, do not develop orchestrators, save your time. You trying to compete with corporations throwing tons of money in their gpus doing the same you are doing . Just relax. Wait till they come with a large carton box and security. Enjoy your last moments in corporate environment which you used to hate most of the time..

Profilbild von Seb
Sebvor 4 Tagen

I can absolutely relate! For the past 2 weeks, I’ve also experimented with a super basic setup: one agent writes the plan, one writes the code (product and tests), one reviews it. Almost no rules. Worked surprisingly well 👍 Except: the architecture was ... basically not existing 🤷‍♂️😉

Profilbild von Pastor Soto
Pastor Sotovor 4 Tagen

What does it mean not considering agents as component of the software design? Great video for me it opens the perspective

Profilbild von Douglas Knesek
Douglas Knesekvor 4 Tagen

Can you please say something more (or have an agent write a list you can share) about what things your harness was orchestrating and what constraints it imposed? That would help us apply what you’ve learned to our harnesses.

Profilbild von Ricky
Rickyvor 4 Tagen

my “harness” is simple when claude finishes a task, a hook invokes codex to review and claude gets the output except for the most simple task - codex always comes back with 2-3 valid findings idk how people gain this one shot confidence

Profilbild von Mariano Blua.
Mariano Blua.vor 4 Tagen

Great topic for a Saturday morning! This is something I've been thinking about a lot too. I wonder how far the single-agent advantage holds, especially on longer tasks. How does context rot factor into that comparison? An architect and a developer, each with a narrow role and focused context, could bring complementary perspectives without one conversation carrying the entire history. Does coordination overhead outweigh those benefits, especially on longer tasks? I'll test it myself. My workflow uses multiple coding agents and models, so I suspect the way responsibilities and context are split matters a lot. I'm also obsessed with autonomous programming. I've been building an IDE for multi-agent coding: define agents through Role.md files, organize them into teams, and run as many team instances as you need, each agent using your preferred coding agent.

Profilbild von Steve
Stevevor 4 Tagen

But I already created the Jira tickets for my team to setup the harnesses 😭

Profilbild von ToolkitSoft®
ToolkitSoft®vor 4 Tagen

This is simplistic. There are many factors and considerations. Take token efficiency, well by shifting work to an external software harness, you can reduce your token consumption. Part of the code review of the LLM output can be checked by the software harness using abstract syntax trees. An external software harness allows you to have an LLM agnostic usage of AI. Lots of benefits.

Profilbild von Everlier
Everliervor 4 Tagen

All behaviors must be composable. All behaviors must be serializable. Assembling an agent should resemble assembling a webpage from ready-made components that you can also rewrite from scratch where needed. I think we're getting there.

Profilbild von a017444
a017444vor 4 Tagen

i don't want to promote my harness so I'm not going to paste a link, but I adapted it from zoocode and it works very well especially in orchestrator mode, and allows for swarms of agents. I definitely couldn't achieve the same with a single agent.

Profilbild von Troy
Troyvor 4 Tagen

Yes! ONE agent holding the entire session context is better. I've been finding that my old protocols, guardrails, and rigid instructions are becoming something I now need to gradually strip away, rather than continue hardening them.

Profilbild von Dust
Dustvor 4 Tagen

It happens again and again, that it turns out to be unwise to invest a lot of effort into building an infrastructure over components that are themselves at the moment improving in unexpected ways. The best tip is just don't touch that stuff now even with a 9 meter GPU rack.

Profilbild von Dushyant Suthar
Dushyant Sutharvor 4 Tagen

Agents start fighting the harness correctness?? and the the rules we try to impose/don't do this that. Collectively pulluting context??I don't know..

Profilbild von Tom Hosiawa
Tom Hosiawavor 4 Tagen

Even though it sounds weird to say it, I still think "habitat" is the better mental framing Agents work in the habitat you provide them. You don't put a harness on them to control them

Profilbild von Reece
Reecevor 4 Tagen

4:20 long Nice

Profilbild von Johannes Ortloff
Johannes Ortloffvor 4 Tagen

Unbelievable pace. Thanks for sharing. How can we best not be absorbed in one idea and missing what will be possible?

Profilbild von Devs
Devsvor 4 Tagen

Moving constraints from deterministic tools to agents worked, but delegating architecture-drift prevention and mutation testing to agents didn't. Agent decides it's unnecessary and skips it, while deterministic tools enforcement gets the result we want.

Profilbild von Jeff Picklyk
Jeff Picklykvor 4 Tagen

For my design, I don't steer the agent's thinking. I gate its transitions. Inside a phase it works freely; to leave, it must write distilled notes a schema demands. Every run ends in a retrospective that proposes config changes. The harness learns, so friction points are fixed.

Profilbild von Tema Jeff Bezos🇬🇭
Tema Jeff Bezos🇬🇭vor 4 Tagen

Grok @elonmusk

Profilbild von Atharva Pandey 💻
Atharva Pandey 💻vor 4 Tagen

q q

Profilbild von Sławek
Sławekvor 4 Tagen

makes me think that orchestration is inherently hard the way multi-threading is and maybe some problems are best done whole by a single agent the way that a single mind can produce a more coherent solution. i suppose a large enough problem that simply is too large for a single mind/agent could be sliced along natural seams (which might be hard to determine a priori) and then those parts done whole by different agents. in other words, slice vertically by feature not horizontally by skill

Profilbild von Craig Glendenning
Craig Glendenningvor 4 Tagen

Engineers engineer. Our humble task now is to get out of our own way.

Profilbild von thewatchinghawk 🇨🇦🇺🇸🇮🇹
thewatchinghawk 🇨🇦🇺🇸🇮🇹vor 4 Tagen

I think breaking down a plan into sub-agents make sense if the task is completely independent. But I agree with your assessment. LLMs are getting better in handling multi task segments of the overall assignment without the need to break it down into smaller subgents.

Profilbild von Matt Parrott
Matt Parrottvor 4 Tagen

With plurnk, I have WORK (new log), FORK (forked log), and BARE (no log or tools, pure inference) as primitives the model can reach for to devise its own arbitrary graphs and topologies as it sees fit. Seems to answer for this question.

Profilbild von Alfred D
Alfred Dvor 4 Tagen

yep… did a very complex harness w langgraph before for my use case - thin harness , fat skills idea from Garry of Y combinator worked very well w my use case

Profilbild von Simon Piscitelli
Simon Piscitellivor 4 Tagen

Maybe the review framework is the problem and CRAP and the rest of the stuff are actually bad

Profilbild von David Vaughn
David Vaughnvor 4 Tagen

@unclebobmartin What do you mean by components in software design?

Profilbild von Malia
Maliavor 4 Tagen

free the agents! ha it depends on the agent but yeah if the task is too small they can't see the bigger picture. they still can't do too big of a task that takes too much context as they'll do stupid things also. so long as the prompt and the task size are in the sweet spot they seem to do well

Profilbild von Nanne
Nannevor 4 Tagen

I keep thinking about Django and Ruby on Rails. Their architectures and design patterns are burned into my brain. It's easy to whiteboard parts of applications and to review sensitive code

Profilbild von Jose Martinez
Jose Martinezvor 4 Tagen

What about harnesses for other engineering disciplines: civil-structural-harnesses? Why only for developers?

Profilbild von Vitalii Ivanov
Vitalii Ivanovvor 4 Tagen

Do you thing your harness would still be useful with pre-SOTA or smaller open weight models that are not that smart?

Profilbild von Tiago ᶜᵃᵐ
Tiago ᶜᵃᵐvor 4 Tagen

I felt the same! I used to decompose tasks very granular, then, it start "feeling" slow, Then i tried the same: "just let the main thread do everything" And then.... WTF?! Anyways... Anybody else is thinking/working on the WHY is this?

Profilbild von Guy
Guyvor 4 Tagen

Be nice to know the prompt/task so we can evaluate too? Got to wonder where this leaves us

Profilbild von Reece
Reecevor 4 Tagen

Question: Have you noticed the harness getting better overtime especially with Grok Bot?

Ähnliche Videos