正在加载视频...
视频加载失败
Continuous self-improvement needs an ever-expanding supply of training environments (goals). SPADE: one model self-plays the Environment Designer and the Reasoning Agent, writing executable, agentic environments that get harder as it improves. Environment scaling on its own. ♠️
146,138 次观看 • 5 天前 •via X (Twitter)
0 条评论
暂无评论
原始帖子的评论将显示在这里
