Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

$0.04 per 1M input tokens 128K context #1 on the Decision Index sub-second decisions Drex is looking less like “another small model” and more like an actual decision layer for agents and LLMs especially if you don’t want to burn a huge model’s tokens just deciding what should happen next

27,308 görüntüleme • 8 gün önce •via X (Twitter)

0 Yorum

Yorum bulunmuyor

Orijinal gönderinin yorumları burada görünecek

Benzer Videolar

If you are trying to understand where AI agents are going, learn harness engineering. A capable model is only one part of an agent system. Once the model begins reading files, calling tools, modifying state and working across many steps, the quality of the system depends increasingly on the software around it. Consider a coding agent working through a large repository. The model can decide that it needs to inspect a file, search for a symbol, make an edit or run a test, but those decisions do not execute themselves. The surrounding runtime has to decide which resources are available, whether the requested action is permitted, how the operation should be performed, what result should be retained, and what information should be presented to the model on the next step. This becomes harder as the run gets longer. As history accumulates, replaying everything can become costly and less effective. The harness has to decide what should remain in context, what should be summarized or retrieved later, and what belongs in persistent state outside the context window. Execution has similar problems. A long-running agent may need to survive an interruption, avoid repeating completed work, enforce permissions around consequential actions, and preserve enough history to reconstruct what happened when the final result is wrong. These are harness problems. The harness is the layer that manages context, tools, execution, state, checkpoints, limits and traces around the model. Harness engineering is the work of designing and improving that layer. Engineers inspect execution traces, evaluate agents on representative tasks, look for recurring failure modes, and then change things such as context selection, tool interfaces, state handling or execution controls. That last part matters because agent failures are often not fixed by changing the model. Sometimes the useful change is in what the model sees, how a tool is exposed, what state is preserved, or what the runtime does after a failed step. As agents take on longer tasks, the demands on this surrounding software grow. Model capability remains essential, but harness engineering is what turns that capability into an execution process that can be controlled, inspected, tested and improved.

Tech with Mak

48,700 görüntüleme • 12 gün önce