Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Next.js installs will soon include version-matched docs, giving agents context on new and recently updated APIs. In our evals, this improved success rates by ~20%. Try it out: 𝚗𝚙𝚡 𝚌𝚛𝚎𝚊𝚝𝚎-𝚗𝚎𝚡𝚝-𝚊𝚙𝚙@𝚌𝚊𝚗𝚊𝚛𝚢

125,088 Aufrufe • vor 5 Monaten •via X (Twitter)

0 Kommentare

Keine Kommentare verfügbar

Kommentare vom Original-Post werden hier angezeigt

Ähnliche Videos

How many AI agents work at your company? We now have over 3,258 agents working alongside 1,300 humans. The crazy part is these agents were created by EVERY EMPLOYEE at our company... sales reps, marketers, customer support, product, eng. Literally EVERYONE. BUT I'm most surprised by the adoption and value that MANAGERS are getting from agents. I used to think that every IC would become a manager of agents. Now I think that managers will very likely manage WAY more agents than their ICs combined. And managers' agents will manage their ICs' agents - overseeing them for human-in-the-loop interactions. When creating agents, we use 100% context from all of your activity, files edited, tasks and projects worked on, hierarchy, skills, and role information. We build a user-based context model to make agents as relatable as possible to the specific human that we're building for. This means they truly understand the nuances of the work and what "great" looks like - because great is very much in the eye of the beholder. Great is by definition, subjective. This is also why the human ENGAGEMENT loops are SO vital to agent value. The iteration AFTER the agent is onboarded is where the MAGIC happens. This is just like a manager managing an IC in real life... you're giving feedback. In this case, though, agents learn INSTANTLY, and they retain the knowledge perfectly and indefinitely. Even though I've been pushing AI for years now to everyone in our company, this was the first time we had truly end-to-end AI adoption and retention. This kind of AI adoption is wild. But the value we're realizing is truly INSANE. Super Agents outnumber our humans nearly 3 to 1. What if you could 3X your workforce overnight? Watch this video to see how 👇

Zeb Evans

425,244 Aufrufe • vor 5 Monaten

Alright, this one’s worth your attention if you’re building or deploying agents. Future AGI just open-sourced their entire platform and i don’t mean a trimmed-down version. this is the full stack: UI, backend, simulation engine, evals, optimization loop, observability, guardrails, gateway, docs. all in one repo. Apache 2.0. I’ve been putting it through its paces on production agents, and what stands out isn’t just the breadth it’s the architecture. Most of the current “agent reliability” stack is fragmented. tracing lives in one tool, evals in another, guardrails somewhere else. you end up manually connecting dots, and the agent itself doesn’t really improve you just keep patching prompts and hoping for the best. This flips that model. It’s built as a closed feedback loop: simulate failures → evaluate in real time → detect production issues → learn from them → generate fixes → validate against real traffic → check regressions → redeploy → monitor again And when something new breaks, the loop just runs again. no manual glue. The simulation piece is especially strong. instead of static test cases, it generates adversarial, multi-turn conversations based on how your agent actually behaves basically hunting for the exact scenarios where your system fails confidently. ran a few thousand simulations on our side… caught things we definitely would’ve missed. Evals run fast (sub-50ms) across modalities. not LLM-as-judge trained classifiers. guardrails are built-in, not layered on top. observability gives you step-level visibility into reasoning, cost, latency, quality. But the real shift is the optimization loop. Most tools tell you *what* broke. this system actually fixes it, validates the fix, and ensures nothing else regresses. That’s the missing layer. It’s clearly built with production in mind not a research demo. and the fact that it’s self-hostable makes it even more relevant if you’re running serious workloads. If you’ve been duct-taping together infra around your agents, this is probably the closest thing to a unified system i’ve seen so far. Worth checking out. If you're serious about deploying reliable AI agents, this is worth a look: 👉 You can also try it instantly (no setup) via their cloud version:

Aakash Verma

22,889 Aufrufe • vor 2 Monaten