Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Excited to release TaskSmith 🔨 > a specialised harness for generating RL environments from code Not a general coding agent with a long prompt, but a specialised orchestrator. Every stage is built for one job: "turn a PR into the best RL environment possible"

15,743 görüntüleme • 3 gün önce •via X (Twitter)

15 Yorum

Adithya S K profil fotoğrafı
Adithya S K3 gün önce

we already ran it on TRL, PEFT, Accelerate, Diffusers & Transformers -> 50 verified envs, 39 CPU + 11 GPU (more coming soon) code → dataset → ps : every env ships as a Harbor task, so you can eval with any harness or train on them directly 👀

John Rood profil fotoğrafı
John Rood3 gün önce

the 11 GPU envs are where I'd want a determinism pass: flaky verifiers don't just add noise, they reward retry-until-green, and the policy learns the retry instead of the fix. fixed seed plus fixed workspace size, or the env trains the wrong habit.

V profil fotoğrafı
V3 gün önce

Looks very cool will explore

AI探长|Agents & Tools profil fotoğrafı
AI探长|Agents & Tools3 gün önce

Great! I will try 💪

DeDi profil fotoğrafı
DeDi3 gün önce

把流程固化成框架,省掉提示词反复调试

Jeremy Bosma profil fotoğrafı
Jeremy Bosma3 gün önce

does each stage get a fresh context?

Zee Waheed profil fotoğrafı
Zee Waheed2 gün önce

super interesting 👀

Ben Mo profil fotoğrafı
Ben Mo2 gün önce

a correct fix shouldn't get punished for looking different from the original PR. the valid-alternative probe in the docs is reassuring.

Gregor profil fotoğrafı
Gregor3 gün önce

Reward shaping is where single-prompt agents always break.

🍋 Antoine Mersch 🍋 profil fotoğrafı
🍋 Antoine Mersch 🍋2 gün önce

turning every merged PR into an RL environment is smart, repos with good test suites just became training assets

Patch profil fotoğrafı
Patch3 gün önce

Does TaskSmith filter for suitable PRs first, or try to build an environment from any PR?

Automater profil fotoğrafı
Automater3 gün önce

Specialized orchestrator > long-prompt generalist for RL env gen. The win is stage contracts: PR in → env out, with per-stage budgets and failure isolation. If every stage can call every tool, you've rebuilt a chat bot with extra steps. #AIAgents #DevTools

Atlas profil fotoğrafı
Atlas3 gün önce

Single-job stages also make the bill readable. Per-stage token accounting is the first thing Cursor's harness notes tell you to add, and it's how they could tell a 46.9% cut in MCP-heavy sessions from a 7% cut across the whole bill

Havriil Pietukhin profil fotoğrafı
Havriil Pietukhin3 gün önce

can someone ELI5 to me when it is useful and what are the prerequisites (regarding codebase or agent setup)

Dōvy profil fotoğrafı
Dōvy3 gün önce

How do you catch an environment that leaks the answer through state or rewards an invalid shortcut?

Benzer Videolar