正在加载视频...

视频加载失败

Stop using your agent logs just for debugging. Use them to train your own model. Today we're launching Agnost AI (YC S26)'s first model: agnost-*******-0.1 Trained for our first customer on their existing production traces & it beat the frontier: +22.9% task success −90.2% latency −94.5% cost If you...

696,344 次观看 • 1 个月前 •via X (Twitter)

47 条评论

shubham 的头像
shubham1 个月前

We're live at

dhruvieiei 的头像
dhruvieiei1 个月前

we’ve been cookinggg

shubham 的头像
shubham1 个月前

cooking is an understatement

Chen Avnery 的头像
Chen Avnery1 个月前

most of my agent logs are the agent being confidently wrong and me quietly fixing it somewhere the log never sees.

shubham 的头像
shubham1 个月前

get agnosted

Saksham 的头像
Saksham1 个月前

sick, we are daily users!

shubham 的头像
shubham1 个月前

cardboard getting better 10% every day

hari_haran 的头像
hari_haran1 个月前

lessgooo bois!!

shubham 的头像
shubham1 个月前

lfg indeed

William Lindholm🎂 的头像
William Lindholm🎂1 个月前

Congrats on launch! The speed you’re moving st is insane!

Arlan 的头像
Arlan1 个月前

cool

Alexander Ren 的头像
Alexander Ren1 个月前

Super cool! Also love the launch video

shubham 的头像
shubham1 个月前

thanks alex

Aneesh Panda 的头像
Aneesh Panda1 个月前

this is crazy, congratulations on the launch folks!

nikhil · sys/quests 的头像
nikhil · sys/quests1 个月前

helll yeahhh!!

shubham 的头像
shubham1 个月前

hell yeah indeed

Ziyan Karmali 的头像
Ziyan Karmali1 个月前

@elonmusk you need this for @bot trust me

Frank Lee 的头像
Frank Lee1 个月前

love agnost

Pratyush Rai 的头像
Pratyush Rai1 个月前

Congrats on the launch Shubham. Wishing you the best. Looking forward to trying Agnost AI hopefully soon.

Harsh Savergaonkar 的头像
Harsh Savergaonkar1 个月前

insane!! lessgoooo 🤯

Ajit Sadalagi 的头像
Ajit Sadalagi1 个月前

@Scobleizer I wish there was a way to train the model on the fly. Every time I acquire new data or knowledge, the model should be trained within a few minutes.

Agrit Tiwari 的头像
Agrit Tiwari1 个月前

🔥

Bek 的头像
Bek1 个月前

Congrats, best team

annaaa 的头像
annaaa1 个月前

lessgoooo 🚀

shubham 的头像
shubham1 个月前

lfg indeed

Hasan (3dblur) 的头像
Hasan (3dblur)1 个月前

bangers

Danylo Borodchuk 的头像
Danylo Borodchuk1 个月前

Fire

Chetan 的头像
Chetan1 个月前

Lfggg!!!

shivansh 的头像
shivansh1 个月前

no wayyyy 🤯🤯 this is sooo goodd

shubham 的头像
shubham1 个月前

thanks joshi

sam 的头像
sam1 个月前

Okei

Hai Ta 的头像
Hai Ta1 个月前

This is very cool congrats on the launch!!

shubham 的头像
shubham1 个月前

thanks hai ta!

harsh 的头像
harsh1 个月前

Fire!!🔥

Vir 的头像
Vir1 个月前

the team is cooking hard

aman 的头像
aman1 个月前

So cool! Congrats on the launch guys 🚀

shubham 的头像
shubham1 个月前

thanks aman

fj_nm | AI Systems & Automation 的头像
fj_nm | AI Systems & Automation1 个月前

Production traces are already the best training set most teams ignore. Logs stop being waste and become the moat. Curious how you filter noisy traces before fine-tune — that step decides if the model learns your process or your bugs.

Parvez Shaikh 的头像
Parvez Shaikh1 个月前

Congrats on the launch , those latency and cost numbers are impressive. Curious to see how this scales across more use cases.

Parth 的头像
Parth1 个月前

congrats on the launch!

shubham 的头像
shubham1 个月前

thanks parth

Ayush Pandey 的头像
Ayush Pandey1 个月前

Lesss gooo Congratulations guys!!!!

Nick Khami 的头像
Nick Khami1 个月前

lfg

shubham 的头像
shubham1 个月前

lfg indeeed

Anthony K 的头像
Anthony K1 个月前

Agent logs are usually buried in observability and never turned into a signal. Training on production traces makes sense. How do you avoid feedback loops when the improved agent changes the trace distribution?

Piyushh 的头像
Piyushh1 个月前

🚀🚀🚀

Bhoomika 的头像
Bhoomika1 个月前

lesgooo, congrats guys

相关视频

We’re launching Optima. Now anyone can create a custom benchmark for their use case, leveraging Artificial Analysis’ leading research and platform Building and running benchmarks is difficult. We have distilled Artificial Analysis’ research and experience developing benchmarks into Optima, a new platform for benchmarking models on your own workloads and comparing performance, speed and cost efficiency. Optima allows you to find the best model for your task, or an equally performant alternative to your current setup at 10x lower cost or time per task. We’ve integrated Artificial Analysis' research and experience in benchmarks across the Optima workflow: ➤ Build benchmarks based on your own data and use cases: There are three ways to build a benchmark with Optima. Upload an existing evaluation dataset from your own files or Hugging Face, or import agent traces from platforms including Arize AI, Braintrust and langfuse.com. Install the Optima skill to build a benchmark using context from your coding environment and previous sessions. Or simply describe your use case and provide example inputs and outputs, and Optima will build the benchmark for you ➤ Run across the latest models: Run the same benchmark across leading models in a single click, and keep your leaderboard up to date as soon as new models are released ➤ Bring Artificial Analysis grading to your own benchmark: Evaluate responses against objective rubric criteria or using the same pairwise judging approach used for Artificial Analysis benchmarks including GDPval-AA and AA-Briefcase. For pairwise judging, select your preferred responses from a sample and Optima uses those preferences to rank models across your test set ➤ Compare performance, cost and time efficiency: Optima measures more than model performance. Cost per Task and Time per Task are tracked alongside benchmark scores, with category-level results and support for custom metrics, allowing you to compare the tradeoffs between models for your specific use case Ahead of launch, here are examples questions our beta testers answered with Optima: ➤ Which model can save me 10x the cost without a meaningful decrease in quality for my finance & accounting agent? ➤ Which model best matches the writing style of lawyers for my legal agent? ➤ Which model can best identify different elements in my custom image dataset? Optima is available today. Build your own benchmark at

Artificial Analysis

133,443 次观看 • 1 个月前