
elune
@elune0x • 1,112 subscribers
growth @kollectivexyz ai influencer | dm open
Videos

Andrej Karpathy spent 8 years at OpenAI and Tesla Last week, he compressed everything he knows into one free 2-hour lecture Agents → Harness → Loops → Graphs → Self-Improving Systems People pay $15K for bootcamps that teach less than this This lecture is better than most paid AI engineering courses You probably don't have 2 hours right now Don't let this disappear from your feed Watch it Then read the guide on harness engineering below
elune216,333 views • 1 month ago

10 agent evals every AI engineer should know 1) golden set a frozen set of cases you run after every prompt model or tool change use it as your baseline to see whether the system improved or quietly broke OpenAI Evals helps you build repeatable benchmark sets and compare model changes → 2) llm as judge a second model scores open ended answers against a written rubric use it when there is no exact output to compare against OpenEvals provides ready made evaluators for LLM applications → 3) rubric scoring score correctness tone safety and cost separately one quality number hides the problem DeepEval helps you create custom metrics and score every dimension independently → 4) trajectory eval grade the path the agent took not only the final answer AgentEvals checks agent actions decisions and tool calls across the full trajectory → 5) tool unit tests test every tool with fixed inputs and outputs no model in the loop MCP Inspector helps you inspect and test MCP servers tools and responses separately → 6) regression suite replay previous runs against every new prompt model or toolset then compare the results Promptfoo helps you run repeatable eval suites catch regressions and add checks to CI → 7) a b in prod split real traffic between two versions and compare actual outcomes GrowthBook provides feature flags controlled experiments and product analytics → 8) human review sample real runs and let a person grade them honestly use it to calibrate your automated judge Argilla helps teams collect human feedback review outputs and build better datasets → 9) shadow run let the candidate run on real traffic while its output is shown to nobody use it before a risky rollout Langfuse helps you trace production runs compare candidates and monitor eval results → 10) red team attack the system before somebody else does jailbreaks prompt injection data leaks and tool abuse Garak scans LLM systems for vulnerabilities and unsafe behavior → offline evals tell you it works online evals tell you it still works you probably do not need all ten today start with the two that would have caught your last outage bookmark this
elune155,026 views • 2 months ago

ANTHROPIC JUST EXPOSED WHY MOST MULTI AGENT CODING SYSTEMS FAIL BEFORE WRITING A SINGLE LINE OF CODE the model is not the problem the plan is one bad dependency cut gets copied across six agents then three builders produce the wrong result in parallel spec → planner → three builders → critic → scribe → PR → back into the spec the planner runs once every agent downstream is forced to inherit that decision builder A owns the API builder B owns the data layer builder C owns the interface none of them reads another lane that is what makes the system fast and what makes a weak plan catastrophic each builder gets its own read write test loop red stays inside the lane green unlocks the critic the critic does not ask whether the code looks good it asks whether the entire plan should exist then the scribe turns the trace into the PR plan hash three diffs test proof critic notes zero summaries written from memory if the PR fails the reason goes back into the spec as a new constraint the next run starts where the last one failed one human step remains approve or send it back five minutes instead of five hours a good plan scales a bad plan multiplies save this before your next agent build ↓
elune34,491 views • 1 month ago

your agent is not the loop the loop is only the smallest layer LOOP vs GRAPH vs HARNESS ENGINEERING most teams still treat all three like one prompt that is why agent failures feel impossible to diagnose LOOP ENGINEERING controls iteration turns retries budgets evaluators exits stalled progress when the run keeps going forever the loop is broken make long-running work survive crashes and restarts → GRAPH ENGINEERING controls structure nodes edges state branches cycles joins checkpoints when the agent takes the wrong path the graph is broken make state and routing explicit → inspect the topology without asking another model to explain it → HARNESS ENGINEERING controls access tools permissions memory sandboxes evals traces humans when the agent touches the wrong system the harness is broken isolate generated code and tool execution from the host → put an eval gate before every model or prompt update → keep one trace across every node model and tool call → without loop engineering it never stops without graph engineering you cannot see the path without harness engineering it can reach anything the prompt lives inside the loop the loop lives inside the graph the graph lives inside the harness the model is the smallest box in the system debug the layer that actually owns the failure save this before the next agent rewrite
elune16,875 views • 2 months ago
No more content to load