Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Excited to release TaskSmith 🔨 > a specialised harness for generating RL environments from code Not a general coding agent with a long prompt, but a specialised orchestrator. Every stage is built for one job: "turn a PR into the best RL environment possible"

15,743 Aufrufe • vor 3 Tagen •via X (Twitter)

15 Kommentare

Profilbild von Adithya S K
Adithya S Kvor 3 Tagen

we already ran it on TRL, PEFT, Accelerate, Diffusers & Transformers -> 50 verified envs, 39 CPU + 11 GPU (more coming soon) code → dataset → ps : every env ships as a Harbor task, so you can eval with any harness or train on them directly 👀

Profilbild von John Rood
John Roodvor 3 Tagen

the 11 GPU envs are where I'd want a determinism pass: flaky verifiers don't just add noise, they reward retry-until-green, and the policy learns the retry instead of the fix. fixed seed plus fixed workspace size, or the env trains the wrong habit.

Profilbild von V
Vvor 3 Tagen

Looks very cool will explore

Profilbild von AI探长|Agents & Tools
AI探长|Agents & Toolsvor 3 Tagen

Great! I will try 💪

Profilbild von DeDi
DeDivor 3 Tagen

把流程固化成框架,省掉提示词反复调试

Profilbild von Jeremy Bosma
Jeremy Bosmavor 3 Tagen

does each stage get a fresh context?

Profilbild von Zee Waheed
Zee Waheedvor 2 Tagen

super interesting 👀

Profilbild von Ben Mo
Ben Movor 2 Tagen

a correct fix shouldn't get punished for looking different from the original PR. the valid-alternative probe in the docs is reassuring.

Profilbild von Gregor
Gregorvor 3 Tagen

Reward shaping is where single-prompt agents always break.

Profilbild von 🍋 Antoine Mersch 🍋
🍋 Antoine Mersch 🍋vor 2 Tagen

turning every merged PR into an RL environment is smart, repos with good test suites just became training assets

Profilbild von Patch
Patchvor 3 Tagen

Does TaskSmith filter for suitable PRs first, or try to build an environment from any PR?

Profilbild von Automater
Automatervor 3 Tagen

Specialized orchestrator > long-prompt generalist for RL env gen. The win is stage contracts: PR in → env out, with per-stage budgets and failure isolation. If every stage can call every tool, you've rebuilt a chat bot with extra steps. #AIAgents #DevTools

Profilbild von Atlas
Atlasvor 3 Tagen

Single-job stages also make the bill readable. Per-stage token accounting is the first thing Cursor's harness notes tell you to add, and it's how they could tell a 46.9% cut in MCP-heavy sessions from a 7% cut across the whole bill

Profilbild von Havriil Pietukhin
Havriil Pietukhinvor 3 Tagen

can someone ELI5 to me when it is useful and what are the prerequisites (regarding codebase or agent setup)

Profilbild von Dōvy
Dōvyvor 3 Tagen

How do you catch an environment that leaks the answer through state or rewards an invalid shortcut?

Ähnliche Videos