Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Stop using your agent logs just for debugging. Use them to train your own model. Today we're launching Agnost AI (YC S26)'s first model: agnost-*******-0.1 Trained for our first customer on their existing production traces & it beat the frontier: +22.9% task success −90.2% latency −94.5% cost If you...

696,344 Aufrufe • vor 1 Monat •via X (Twitter)

47 Kommentare

Profilbild von shubham
shubhamvor 1 Monat

We're live at

Profilbild von dhruvieiei
dhruvieieivor 1 Monat

we’ve been cookinggg

Profilbild von shubham
shubhamvor 1 Monat

cooking is an understatement

Profilbild von Chen Avnery
Chen Avneryvor 1 Monat

most of my agent logs are the agent being confidently wrong and me quietly fixing it somewhere the log never sees.

Profilbild von shubham
shubhamvor 1 Monat

get agnosted

Profilbild von Saksham
Sakshamvor 1 Monat

sick, we are daily users!

Profilbild von shubham
shubhamvor 1 Monat

cardboard getting better 10% every day

Profilbild von hari_haran
hari_haranvor 1 Monat

lessgooo bois!!

Profilbild von shubham
shubhamvor 1 Monat

lfg indeed

Profilbild von William Lindholm🎂
William Lindholm🎂vor 1 Monat

Congrats on launch! The speed you’re moving st is insane!

Profilbild von Arlan
Arlanvor 1 Monat

cool

Profilbild von Alexander Ren
Alexander Renvor 1 Monat

Super cool! Also love the launch video

Profilbild von shubham
shubhamvor 1 Monat

thanks alex

Profilbild von Aneesh Panda
Aneesh Pandavor 1 Monat

this is crazy, congratulations on the launch folks!

Profilbild von nikhil · sys/quests
nikhil · sys/questsvor 1 Monat

helll yeahhh!!

Profilbild von shubham
shubhamvor 1 Monat

hell yeah indeed

Profilbild von Ziyan Karmali
Ziyan Karmalivor 1 Monat

@elonmusk you need this for @bot trust me

Profilbild von Frank Lee
Frank Leevor 1 Monat

love agnost

Profilbild von Pratyush Rai
Pratyush Raivor 1 Monat

Congrats on the launch Shubham. Wishing you the best. Looking forward to trying Agnost AI hopefully soon.

Profilbild von Harsh Savergaonkar
Harsh Savergaonkarvor 1 Monat

insane!! lessgoooo 🤯

Profilbild von Ajit Sadalagi
Ajit Sadalagivor 1 Monat

@Scobleizer I wish there was a way to train the model on the fly. Every time I acquire new data or knowledge, the model should be trained within a few minutes.

Profilbild von Agrit Tiwari
Agrit Tiwarivor 1 Monat

🔥

Profilbild von Bek
Bekvor 1 Monat

Congrats, best team

Profilbild von annaaa
annaaavor 1 Monat

lessgoooo 🚀

Profilbild von shubham
shubhamvor 1 Monat

lfg indeed

Profilbild von Hasan (3dblur)
Hasan (3dblur)vor 1 Monat

bangers

Profilbild von Danylo Borodchuk
Danylo Borodchukvor 1 Monat

Fire

Profilbild von Chetan
Chetanvor 1 Monat

Lfggg!!!

Profilbild von shivansh
shivanshvor 1 Monat

no wayyyy 🤯🤯 this is sooo goodd

Profilbild von shubham
shubhamvor 1 Monat

thanks joshi

Profilbild von sam
samvor 1 Monat

Okei

Profilbild von Hai Ta
Hai Tavor 1 Monat

This is very cool congrats on the launch!!

Profilbild von shubham
shubhamvor 1 Monat

thanks hai ta!

Profilbild von harsh
harshvor 1 Monat

Fire!!🔥

Profilbild von Vir
Virvor 1 Monat

the team is cooking hard

Profilbild von aman
amanvor 1 Monat

So cool! Congrats on the launch guys 🚀

Profilbild von shubham
shubhamvor 1 Monat

thanks aman

Profilbild von fj_nm | AI Systems & Automation
fj_nm | AI Systems & Automationvor 1 Monat

Production traces are already the best training set most teams ignore. Logs stop being waste and become the moat. Curious how you filter noisy traces before fine-tune — that step decides if the model learns your process or your bugs.

Profilbild von Parvez Shaikh
Parvez Shaikhvor 1 Monat

Congrats on the launch , those latency and cost numbers are impressive. Curious to see how this scales across more use cases.

Profilbild von Parth
Parthvor 1 Monat

congrats on the launch!

Profilbild von shubham
shubhamvor 1 Monat

thanks parth

Profilbild von Ayush Pandey
Ayush Pandeyvor 1 Monat

Lesss gooo Congratulations guys!!!!

Profilbild von Nick Khami
Nick Khamivor 1 Monat

lfg

Profilbild von shubham
shubhamvor 1 Monat

lfg indeeed

Profilbild von Anthony K
Anthony Kvor 1 Monat

Agent logs are usually buried in observability and never turned into a signal. Training on production traces makes sense. How do you avoid feedback loops when the improved agent changes the trace distribution?

Profilbild von Piyushh
Piyushhvor 1 Monat

🚀🚀🚀

Profilbild von Bhoomika
Bhoomikavor 1 Monat

lesgooo, congrats guys

Ähnliche Videos

We’re launching Optima. Now anyone can create a custom benchmark for their use case, leveraging Artificial Analysis’ leading research and platform Building and running benchmarks is difficult. We have distilled Artificial Analysis’ research and experience developing benchmarks into Optima, a new platform for benchmarking models on your own workloads and comparing performance, speed and cost efficiency. Optima allows you to find the best model for your task, or an equally performant alternative to your current setup at 10x lower cost or time per task. We’ve integrated Artificial Analysis' research and experience in benchmarks across the Optima workflow: ➤ Build benchmarks based on your own data and use cases: There are three ways to build a benchmark with Optima. Upload an existing evaluation dataset from your own files or Hugging Face, or import agent traces from platforms including Arize AI, Braintrust and langfuse.com. Install the Optima skill to build a benchmark using context from your coding environment and previous sessions. Or simply describe your use case and provide example inputs and outputs, and Optima will build the benchmark for you ➤ Run across the latest models: Run the same benchmark across leading models in a single click, and keep your leaderboard up to date as soon as new models are released ➤ Bring Artificial Analysis grading to your own benchmark: Evaluate responses against objective rubric criteria or using the same pairwise judging approach used for Artificial Analysis benchmarks including GDPval-AA and AA-Briefcase. For pairwise judging, select your preferred responses from a sample and Optima uses those preferences to rank models across your test set ➤ Compare performance, cost and time efficiency: Optima measures more than model performance. Cost per Task and Time per Task are tracked alongside benchmark scores, with category-level results and support for custom metrics, allowing you to compare the tradeoffs between models for your specific use case Ahead of launch, here are examples questions our beta testers answered with Optima: ➤ Which model can save me 10x the cost without a meaningful decrease in quality for my finance & accounting agent? ➤ Which model best matches the writing style of lawyers for my legal agent? ➤ Which model can best identify different elements in my custom image dataset? Optima is available today. Build your own benchmark at

Artificial Analysis

133,443 Aufrufe • vor 1 Monat