正在加载视频...

视频加载失败

Excited to release TaskSmith 🔨 > a specialised harness for generating RL environments from code Not a general coding agent with a long prompt, but a specialised orchestrator. Every stage is built for one job: "turn a PR into the best RL environment possible"

15,743 次观看 • 3 天前 •via X (Twitter)

15 条评论

Adithya S K 的头像
Adithya S K3 天前

we already ran it on TRL, PEFT, Accelerate, Diffusers & Transformers -> 50 verified envs, 39 CPU + 11 GPU (more coming soon) code → dataset → ps : every env ships as a Harbor task, so you can eval with any harness or train on them directly 👀

John Rood 的头像
John Rood3 天前

the 11 GPU envs are where I'd want a determinism pass: flaky verifiers don't just add noise, they reward retry-until-green, and the policy learns the retry instead of the fix. fixed seed plus fixed workspace size, or the env trains the wrong habit.

V 的头像
V3 天前

Looks very cool will explore

AI探长|Agents & Tools 的头像
AI探长|Agents & Tools3 天前

Great! I will try 💪

DeDi 的头像
DeDi3 天前

把流程固化成框架,省掉提示词反复调试

Jeremy Bosma 的头像
Jeremy Bosma3 天前

does each stage get a fresh context?

Zee Waheed 的头像
Zee Waheed2 天前

super interesting 👀

Ben Mo 的头像
Ben Mo2 天前

a correct fix shouldn't get punished for looking different from the original PR. the valid-alternative probe in the docs is reassuring.

Gregor 的头像
Gregor3 天前

Reward shaping is where single-prompt agents always break.

🍋 Antoine Mersch 🍋 的头像
🍋 Antoine Mersch 🍋2 天前

turning every merged PR into an RL environment is smart, repos with good test suites just became training assets

Patch 的头像
Patch3 天前

Does TaskSmith filter for suitable PRs first, or try to build an environment from any PR?

Automater 的头像
Automater3 天前

Specialized orchestrator > long-prompt generalist for RL env gen. The win is stage contracts: PR in → env out, with per-stage budgets and failure isolation. If every stage can call every tool, you've rebuilt a chat bot with extra steps. #AIAgents #DevTools

Atlas 的头像
Atlas3 天前

Single-job stages also make the bill readable. Per-stage token accounting is the first thing Cursor's harness notes tell you to add, and it's how they could tell a 46.9% cut in MCP-heavy sessions from a 7% cut across the whole bill

Havriil Pietukhin 的头像
Havriil Pietukhin3 天前

can someone ELI5 to me when it is useful and what are the prerequisites (regarding codebase or agent setup)

Dōvy 的头像
Dōvy3 天前

How do you catch an environment that leaks the answer through state or rewards an invalid shortcut?

相关视频