Loading video...
Video Failed to Load
Excited to release TaskSmith 🔨 > a specialised harness for generating RL environments from code Not a general coding agent with a long prompt, but a specialised orchestrator. Every stage is built for one job: "turn a PR into the best RL environment possible"
15,743 views • 3 days ago •via X (Twitter)
15 Comments

we already ran it on TRL, PEFT, Accelerate, Diffusers & Transformers -> 50 verified envs, 39 CPU + 11 GPU (more coming soon) code → dataset → ps : every env ships as a Harbor task, so you can eval with any harness or train on them directly 👀

the 11 GPU envs are where I'd want a determinism pass: flaky verifiers don't just add noise, they reward retry-until-green, and the policy learns the retry instead of the fix. fixed seed plus fixed workspace size, or the env trains the wrong habit.

Looks very cool will explore

Great! I will try 💪

把流程固化成框架,省掉提示词反复调试

does each stage get a fresh context?

super interesting 👀

a correct fix shouldn't get punished for looking different from the original PR. the valid-alternative probe in the docs is reassuring.

Reward shaping is where single-prompt agents always break.

turning every merged PR into an RL environment is smart, repos with good test suites just became training assets

Does TaskSmith filter for suitable PRs first, or try to build an environment from any PR?

Specialized orchestrator > long-prompt generalist for RL env gen. The win is stage contracts: PR in → env out, with per-stage budgets and failure isolation. If every stage can call every tool, you've rebuilt a chat bot with extra steps. #AIAgents #DevTools

Single-job stages also make the bill readable. Per-stage token accounting is the first thing Cursor's harness notes tell you to add, and it's how they could tell a 46.9% cut in MCP-heavy sessions from a 7% cut across the whole bill

can someone ELI5 to me when it is useful and what are the prerequisites (regarding codebase or agent setup)

How do you catch an environment that leaks the answer through state or rewards an invalid shortcut?
