Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Stop using your agent logs just for debugging. Use them to train your own model. Today we're launching Agnost AI (YC S26)'s first model: agnost-*******-0.1 Trained for our first customer on their existing production traces & it beat the frontier: +22.9% task success −90.2% latency −94.5% cost If you...

696,344 görüntüleme • 1 ay önce •via X (Twitter)

47 Yorum

shubham profil fotoğrafı
shubham1 ay önce

We're live at

dhruvieiei profil fotoğrafı
dhruvieiei1 ay önce

we’ve been cookinggg

shubham profil fotoğrafı
shubham1 ay önce

cooking is an understatement

Chen Avnery profil fotoğrafı
Chen Avnery1 ay önce

most of my agent logs are the agent being confidently wrong and me quietly fixing it somewhere the log never sees.

shubham profil fotoğrafı
shubham1 ay önce

get agnosted

Saksham profil fotoğrafı
Saksham1 ay önce

sick, we are daily users!

shubham profil fotoğrafı
shubham1 ay önce

cardboard getting better 10% every day

hari_haran profil fotoğrafı
hari_haran1 ay önce

lessgooo bois!!

shubham profil fotoğrafı
shubham1 ay önce

lfg indeed

William Lindholm🎂 profil fotoğrafı
William Lindholm🎂1 ay önce

Congrats on launch! The speed you’re moving st is insane!

Arlan profil fotoğrafı
Arlan1 ay önce

cool

Alexander Ren profil fotoğrafı
Alexander Ren1 ay önce

Super cool! Also love the launch video

shubham profil fotoğrafı
shubham1 ay önce

thanks alex

Aneesh Panda profil fotoğrafı
Aneesh Panda1 ay önce

this is crazy, congratulations on the launch folks!

nikhil · sys/quests profil fotoğrafı
nikhil · sys/quests1 ay önce

helll yeahhh!!

shubham profil fotoğrafı
shubham1 ay önce

hell yeah indeed

Ziyan Karmali profil fotoğrafı
Ziyan Karmali1 ay önce

@elonmusk you need this for @bot trust me

Frank Lee profil fotoğrafı
Frank Lee1 ay önce

love agnost

Pratyush Rai profil fotoğrafı
Pratyush Rai1 ay önce

Congrats on the launch Shubham. Wishing you the best. Looking forward to trying Agnost AI hopefully soon.

Harsh Savergaonkar profil fotoğrafı
Harsh Savergaonkar1 ay önce

insane!! lessgoooo 🤯

Ajit Sadalagi profil fotoğrafı
Ajit Sadalagi1 ay önce

@Scobleizer I wish there was a way to train the model on the fly. Every time I acquire new data or knowledge, the model should be trained within a few minutes.

Agrit Tiwari profil fotoğrafı
Agrit Tiwari1 ay önce

🔥

Bek profil fotoğrafı
Bek1 ay önce

Congrats, best team

annaaa profil fotoğrafı
annaaa1 ay önce

lessgoooo 🚀

shubham profil fotoğrafı
shubham1 ay önce

lfg indeed

Hasan (3dblur) profil fotoğrafı
Hasan (3dblur)1 ay önce

bangers

Danylo Borodchuk profil fotoğrafı
Danylo Borodchuk1 ay önce

Fire

Chetan profil fotoğrafı
Chetan1 ay önce

Lfggg!!!

shivansh profil fotoğrafı
shivansh1 ay önce

no wayyyy 🤯🤯 this is sooo goodd

shubham profil fotoğrafı
shubham1 ay önce

thanks joshi

sam profil fotoğrafı
sam1 ay önce

Okei

Hai Ta profil fotoğrafı
Hai Ta1 ay önce

This is very cool congrats on the launch!!

shubham profil fotoğrafı
shubham1 ay önce

thanks hai ta!

harsh profil fotoğrafı
harsh1 ay önce

Fire!!🔥

Vir profil fotoğrafı
Vir1 ay önce

the team is cooking hard

aman profil fotoğrafı
aman1 ay önce

So cool! Congrats on the launch guys 🚀

shubham profil fotoğrafı
shubham1 ay önce

thanks aman

fj_nm | AI Systems & Automation profil fotoğrafı
fj_nm | AI Systems & Automation1 ay önce

Production traces are already the best training set most teams ignore. Logs stop being waste and become the moat. Curious how you filter noisy traces before fine-tune — that step decides if the model learns your process or your bugs.

Parvez Shaikh profil fotoğrafı
Parvez Shaikh1 ay önce

Congrats on the launch , those latency and cost numbers are impressive. Curious to see how this scales across more use cases.

Parth profil fotoğrafı
Parth1 ay önce

congrats on the launch!

shubham profil fotoğrafı
shubham1 ay önce

thanks parth

Ayush Pandey profil fotoğrafı
Ayush Pandey1 ay önce

Lesss gooo Congratulations guys!!!!

Nick Khami profil fotoğrafı
Nick Khami1 ay önce

lfg

shubham profil fotoğrafı
shubham1 ay önce

lfg indeeed

Anthony K profil fotoğrafı
Anthony K1 ay önce

Agent logs are usually buried in observability and never turned into a signal. Training on production traces makes sense. How do you avoid feedback loops when the improved agent changes the trace distribution?

Piyushh profil fotoğrafı
Piyushh1 ay önce

🚀🚀🚀

Bhoomika profil fotoğrafı
Bhoomika1 ay önce

lesgooo, congrats guys

Benzer Videolar

We’re launching Optima. Now anyone can create a custom benchmark for their use case, leveraging Artificial Analysis’ leading research and platform Building and running benchmarks is difficult. We have distilled Artificial Analysis’ research and experience developing benchmarks into Optima, a new platform for benchmarking models on your own workloads and comparing performance, speed and cost efficiency. Optima allows you to find the best model for your task, or an equally performant alternative to your current setup at 10x lower cost or time per task. We’ve integrated Artificial Analysis' research and experience in benchmarks across the Optima workflow: ➤ Build benchmarks based on your own data and use cases: There are three ways to build a benchmark with Optima. Upload an existing evaluation dataset from your own files or Hugging Face, or import agent traces from platforms including Arize AI, Braintrust and langfuse.com. Install the Optima skill to build a benchmark using context from your coding environment and previous sessions. Or simply describe your use case and provide example inputs and outputs, and Optima will build the benchmark for you ➤ Run across the latest models: Run the same benchmark across leading models in a single click, and keep your leaderboard up to date as soon as new models are released ➤ Bring Artificial Analysis grading to your own benchmark: Evaluate responses against objective rubric criteria or using the same pairwise judging approach used for Artificial Analysis benchmarks including GDPval-AA and AA-Briefcase. For pairwise judging, select your preferred responses from a sample and Optima uses those preferences to rank models across your test set ➤ Compare performance, cost and time efficiency: Optima measures more than model performance. Cost per Task and Time per Task are tracked alongside benchmark scores, with category-level results and support for custom metrics, allowing you to compare the tradeoffs between models for your specific use case Ahead of launch, here are examples questions our beta testers answered with Optima: ➤ Which model can save me 10x the cost without a meaningful decrease in quality for my finance & accounting agent? ➤ Which model best matches the writing style of lawyers for my legal agent? ➤ Which model can best identify different elements in my custom image dataset? Optima is available today. Build your own benchmark at

Artificial Analysis

133,443 görüntüleme • 1 ay önce