Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Anthropic engineer: "Harness engineering is a really rare skill right now - and it's hard to know if you're even good at it." how he actually thinks about building the full system: • step 1 → the model. already smarter than we use it. "freeze development - you'd find...

44,590 Aufrufe • vor 2 Monaten •via X (Twitter)

23 Kommentare

Profilbild von Gipp 🦅
Gipp 🦅vor 2 Monaten

how does deleting steering impact real-world agent reliability?

Profilbild von Codez
Codezvor 2 Monaten

Delete steering, and reliability drops fast - you’re not simplifying the agent, you’re removing the thing that keeps autonomy honest.

Profilbild von Granite
Granitevor 2 Monaten

still amazed by how smart these engineers sound btw i'm actively studying harness engineering rn

Profilbild von Codez
Codezvor 2 Monaten

Also, I’m glad we get to learn directly from the Anthropic team itself. Yeah, harness engineering is where I’m spending most of my time right now.

Profilbild von rewind
rewindvor 2 Monaten

model was never the bottleneck

Profilbild von Codez
Codezvor 2 Monaten

Maybe in the beginning of AI, but definitely not now.

Profilbild von Yarchi
Yarchivor 2 Monaten

Very clear explanation, thanks for sharing

Profilbild von AdiiX
AdiiXvor 2 Monaten

the whole point is step 1 is done, everyone still tuning prompts is fighting last years bottleneck

Profilbild von Codez
Codezvor 2 Monaten

yeah, sadly true

Profilbild von 0xRafy
0xRafyvor 2 Monaten

Thariq is always shipping alpha, great share bro !

Profilbild von Codez
Codezvor 2 Monaten

Yeah, a real legend. I really loved the system he shared on the podcast.

Profilbild von Fabrizio Degeno
Fabrizio Degenovor 2 Monaten

Something new in my feed

Profilbild von Codez
Codezvor 2 Monaten

yeah, podcast was posted less then 24-h ago

Profilbild von Muhammad Ali
Muhammad Alivor 2 Monaten

The harness insight is underrated but "delete the steering" only works when your codebase is small enough to fit in context. Past 50K lines the model needs explicit guidance or it drifts into fixing the wrong file. What repo size are you testing this approach on?

Profilbild von Genius💡💹🧲 🤖
Genius💡💹🧲 🤖vor 2 Monaten

Learning directly from builders always gives way better insights than random takes

Profilbild von Electrik Dreams
Electrik Dreamsvor 2 Monaten

RAG as anti-pattern will get pushback, but it's technically sound. At 4k tokens, vector retrieval was necessary infrastructure. At 200k, you're adding latency and fragility for diminishing returns. grep scales.

Profilbild von Adam
Adamvor 2 Monaten

The distinction between a stronger model and a better harness is the part most builders miss. Your long-form AI articles deserve one clear home where readers can browse the roadmap, newsletter and best systems—not just a Substack profile. Would a quick homepage concept be useful?

Profilbild von FenixFlow
FenixFlowvor 2 Monaten

Deleting steering cuts down on prompt brittleness and over-constraint, letting the model use its native reasoning more freely. In real workflows we've seen it boost reliability on multi-step tasks (fewer hallucinations from forced plans), but it only works well with solid agent harnesses + clean context to catch drift. The engineer basically said they stopped babysitting and started trusting the model more - results improved.

Profilbild von 0xSlyth
0xSlythvor 2 Monaten

skip the hype build the system

Profilbild von Miles S.
Miles S.vor 2 Monaten

i’m trying grep before touching another vector database

Profilbild von Neo
Neovor 2 Monaten

The real insight isn't any single step. it's that harness quality is unmeasurable by most teams. You can A/B test a model. Almost nobody has a rigorous way to A/B test a harness, which is exactly why it stays a rare skill.

Profilbild von whydeso
whydesovor 2 Monaten

Optimizing all four steps instead of just the model is key

Profilbild von Shoopy
Shoopyvor 2 Monaten

"RAG is an anti-pattern, use grep" is such a spicy line lol. Booking this one

Ähnliche Videos