Загрузка видео...
Не удалось загрузить видео
Orchard is an open-source framework for the research community to train and evaluate AI agents across task types. It reduces complexity while supporting strong performance from smaller models by enabling researchers to reuse the same infrastructure.
26,103 просмотров • 1 месяц назад •via X (Twitter)
Комментарии: 11

🌳 Orchard is now fully open—code, data & new results! Give Orchard a try! 🚀 📄 Paper: 📝 Blog: 💻 Code: 📦 Data:

Reusable eval infra is the piece everyone rebuilds from scratch, good to see it opened up. The gap I keep running into is multi session evaluation. Most agent benchmarks score a single run, but the failures that actually hurt in production show up when the agent needs something it learned two sessions ago. Do any of the Orchard task types carry state between runs?

Reusable infrastructure can make agent research easier to compare across task types. It would be valuable to publish recovery behavior alongside task success.

Interesting release

PaRDeS

Reducing complexity while still supporting smaller models is a good balance.

Reusable infrastructure for training and evaluation can make smaller models far more competitive. I would love to see cost per verified task across model sizes on the same Orchard workflow.

Reusing the same infrastructure across task types is smart design

This looks amazing! Open-source frameworks like this are game-changers for the research community. Excited to see how this streamlines model evaluation across different tasks!

Shared infrastructure is becoming just as important as better models.

开源这点挺关键,复现实验能少踩不少坑

