Загрузка видео...

Не удалось загрузить видео

На главную

Stop using your agent logs just for debugging. Use them to train your own model. Today we're launching Agnost AI (YC S26)'s first model: agnost-*******-0.1 Trained for our first customer on their existing production traces & it beat the frontier: +22.9% task success −90.2% latency −94.5% cost If you...

696,344 просмотров • 1 месяц назад •via X (Twitter)

Комментарии: 47

Фото профиля shubham
shubham1 месяц назад

We're live at

Фото профиля dhruvieiei
dhruvieiei1 месяц назад

we’ve been cookinggg

Фото профиля shubham
shubham1 месяц назад

cooking is an understatement

Фото профиля Chen Avnery
Chen Avnery1 месяц назад

most of my agent logs are the agent being confidently wrong and me quietly fixing it somewhere the log never sees.

Фото профиля shubham
shubham1 месяц назад

get agnosted

Фото профиля Saksham
Saksham1 месяц назад

sick, we are daily users!

Фото профиля shubham
shubham1 месяц назад

cardboard getting better 10% every day

Фото профиля hari_haran
hari_haran1 месяц назад

lessgooo bois!!

Фото профиля shubham
shubham1 месяц назад

lfg indeed

Фото профиля William Lindholm🎂
William Lindholm🎂1 месяц назад

Congrats on launch! The speed you’re moving st is insane!

Фото профиля Arlan
Arlan1 месяц назад

cool

Фото профиля Alexander Ren
Alexander Ren1 месяц назад

Super cool! Also love the launch video

Фото профиля shubham
shubham1 месяц назад

thanks alex

Фото профиля Aneesh Panda
Aneesh Panda1 месяц назад

this is crazy, congratulations on the launch folks!

Фото профиля nikhil · sys/quests
nikhil · sys/quests1 месяц назад

helll yeahhh!!

Фото профиля shubham
shubham1 месяц назад

hell yeah indeed

Фото профиля Ziyan Karmali
Ziyan Karmali1 месяц назад

@elonmusk you need this for @bot trust me

Фото профиля Frank Lee
Frank Lee1 месяц назад

love agnost

Фото профиля Pratyush Rai
Pratyush Rai1 месяц назад

Congrats on the launch Shubham. Wishing you the best. Looking forward to trying Agnost AI hopefully soon.

Фото профиля Harsh Savergaonkar
Harsh Savergaonkar1 месяц назад

insane!! lessgoooo 🤯

Фото профиля Ajit Sadalagi
Ajit Sadalagi1 месяц назад

@Scobleizer I wish there was a way to train the model on the fly. Every time I acquire new data or knowledge, the model should be trained within a few minutes.

Фото профиля Agrit Tiwari
Agrit Tiwari1 месяц назад

🔥

Фото профиля Bek
Bek1 месяц назад

Congrats, best team

Фото профиля annaaa
annaaa1 месяц назад

lessgoooo 🚀

Фото профиля shubham
shubham1 месяц назад

lfg indeed

Фото профиля Hasan (3dblur)
Hasan (3dblur)1 месяц назад

bangers

Фото профиля Danylo Borodchuk
Danylo Borodchuk1 месяц назад

Fire

Фото профиля Chetan
Chetan1 месяц назад

Lfggg!!!

Фото профиля shivansh
shivansh1 месяц назад

no wayyyy 🤯🤯 this is sooo goodd

Фото профиля shubham
shubham1 месяц назад

thanks joshi

Фото профиля sam
sam1 месяц назад

Okei

Фото профиля Hai Ta
Hai Ta1 месяц назад

This is very cool congrats on the launch!!

Фото профиля shubham
shubham1 месяц назад

thanks hai ta!

Фото профиля harsh
harsh1 месяц назад

Fire!!🔥

Фото профиля Vir
Vir1 месяц назад

the team is cooking hard

Фото профиля aman
aman1 месяц назад

So cool! Congrats on the launch guys 🚀

Фото профиля shubham
shubham1 месяц назад

thanks aman

Фото профиля fj_nm | AI Systems & Automation
fj_nm | AI Systems & Automation1 месяц назад

Production traces are already the best training set most teams ignore. Logs stop being waste and become the moat. Curious how you filter noisy traces before fine-tune — that step decides if the model learns your process or your bugs.

Фото профиля Parvez Shaikh
Parvez Shaikh1 месяц назад

Congrats on the launch , those latency and cost numbers are impressive. Curious to see how this scales across more use cases.

Фото профиля Parth
Parth1 месяц назад

congrats on the launch!

Фото профиля shubham
shubham1 месяц назад

thanks parth

Фото профиля Ayush Pandey
Ayush Pandey1 месяц назад

Lesss gooo Congratulations guys!!!!

Фото профиля Nick Khami
Nick Khami1 месяц назад

lfg

Фото профиля shubham
shubham1 месяц назад

lfg indeeed

Фото профиля Anthony K
Anthony K1 месяц назад

Agent logs are usually buried in observability and never turned into a signal. Training on production traces makes sense. How do you avoid feedback loops when the improved agent changes the trace distribution?

Фото профиля Piyushh
Piyushh1 месяц назад

🚀🚀🚀

Фото профиля Bhoomika
Bhoomika1 месяц назад

lesgooo, congrats guys

Похожие видео

We’re launching Optima. Now anyone can create a custom benchmark for their use case, leveraging Artificial Analysis’ leading research and platform Building and running benchmarks is difficult. We have distilled Artificial Analysis’ research and experience developing benchmarks into Optima, a new platform for benchmarking models on your own workloads and comparing performance, speed and cost efficiency. Optima allows you to find the best model for your task, or an equally performant alternative to your current setup at 10x lower cost or time per task. We’ve integrated Artificial Analysis' research and experience in benchmarks across the Optima workflow: ➤ Build benchmarks based on your own data and use cases: There are three ways to build a benchmark with Optima. Upload an existing evaluation dataset from your own files or Hugging Face, or import agent traces from platforms including Arize AI, Braintrust and langfuse.com. Install the Optima skill to build a benchmark using context from your coding environment and previous sessions. Or simply describe your use case and provide example inputs and outputs, and Optima will build the benchmark for you ➤ Run across the latest models: Run the same benchmark across leading models in a single click, and keep your leaderboard up to date as soon as new models are released ➤ Bring Artificial Analysis grading to your own benchmark: Evaluate responses against objective rubric criteria or using the same pairwise judging approach used for Artificial Analysis benchmarks including GDPval-AA and AA-Briefcase. For pairwise judging, select your preferred responses from a sample and Optima uses those preferences to rank models across your test set ➤ Compare performance, cost and time efficiency: Optima measures more than model performance. Cost per Task and Time per Task are tracked alongside benchmark scores, with category-level results and support for custom metrics, allowing you to compare the tradeoffs between models for your specific use case Ahead of launch, here are examples questions our beta testers answered with Optima: ➤ Which model can save me 10x the cost without a meaningful decrease in quality for my finance & accounting agent? ➤ Which model best matches the writing style of lawyers for my legal agent? ➤ Which model can best identify different elements in my custom image dataset? Optima is available today. Build your own benchmark at

Artificial Analysis

133,443 просмотров • 1 месяц назад