
Hamza Tahir
@htahir111 • 1,138 subscribers
Co-founder @ ZenML. Building Kitaru: turn agent traces to reliable evals.
Videos

Well this went viral. Seems like we need a better way to visualize agentic traces, especially for people who are not crazy technical enough to diagnose a JSON trace. Here's one UX we're exploring in Kitaru ( - a Tinder-style approach to evaluating your agent traces. You can swipe left and right (and up), and leave notes on what went wrong with a mic. There are so many UX innovations that have not happened yet. Leave more ideas in the comments if you have any 👇
Hamza Tahir16,184 Aufrufe • vor 16 Tagen

You probably have thousands of agent traces sitting in production. And most of them are doing nothing for you. Today we’re launching the new Kitaru to change that. Kitaru ingests traces from Langfuse, Braintrust, or any OTel source, investigates what went right and wrong, groups recurring failures into cohorts, and builds evaluators around them. Then you can replay those same production cases against a different model, prompt, context, or agent setup and ask: What would have happened if we changed this? Instead of traces being something you inspect after a failure, they become part of a continuous improvement loop. Open source. Free to use. Try it out today:
Hamza Tahir22,264 Aufrufe • vor 1 Monat
Keine weiteren Inhalte verfügbar