ๆญฃๅœจๅŠ ่ฝฝ่ง†้ข‘...

่ง†้ข‘ๅŠ ่ฝฝๅคฑ่ดฅ

just shipped taskmaster v0.15 ๐Ÿš€ โ†’ add-task is now context-aware ๐Ÿ”ฅ โ†’ research flag for parse-prd โ†’ new move command โ†’ analyze-complexity of specific tasks โ†’ add tasks without a PRD โ†’ ollama improved โ†’ celebrating 10k โญ๏ธ on GitHub +++ follow + bookmark + dive in ๐Ÿ‘€๐Ÿ‘‡

15,474 ๆฌก่ง‚็œ‹ โ€ข 1 ๅนดๅ‰ โ€ขvia X (Twitter)

0 ๆก่ฏ„่ฎบ

ๆš‚ๆ— ่ฏ„่ฎบ

ๅŽŸๅง‹ๅธ–ๅญ็š„่ฏ„่ฎบๅฐ†ๆ˜พ็คบๅœจ่ฟ™้‡Œ

็›ธๅ…ณ่ง†้ข‘

Super excited to share ๐Ÿง MLGym ๐Ÿฆพ โ€“ the first Gym environment for AI Research Agents ๐Ÿค–๐Ÿ”ฌ We introduce MLGym and MLGym-Bench, a new framework and benchmark for evaluating and developing LLM agents on AI research tasks. The key contributions of our work are: ๐Ÿ•น๏ธ Enables the exploration of different training algorithms for AI Research Agents such as RL ๐Ÿ› ๏ธ Provides a flexible evaluation framework that can accommodate different artifacts such as models, algorithms, or predictions ๐Ÿค– Allows researchers to evaluate any model without the need to develop a custom agentic harness ๐ŸŽฏ Introduces 13 diverse open-ended AI Research tasks for evaluating AI Research Agents on a wide range of domains such as computer vision, natural language processing, reinforcement learning, game theory, and logical reasoning. ๐Ÿ“ˆ Proposes a new evaluation metric for AI Research Agents MLGym makes it easy to: 1) Add new tasks 2) Evaluate new models 3) Integrate new agents Check out a video of the MLGym Agent to see how it performs the full pipeline of idea generation๐Ÿ’ก, implementation ๐Ÿ‘ฉโ€๐Ÿ’ป, experimentation ๐Ÿ‘ฉโ€๐Ÿ”ฌ, and iteration ๐Ÿ”„ to improve on ML tasks. Huge thanks to the exceptionally talented Deepak Nathani who led this work and to all the other amazing collaborators who made this possible ๐Ÿ™๐Ÿซถ๐Ÿš€

Roberta Raileanu

105,200 ๆฌก่ง‚็œ‹ โ€ข 1 ๅนดๅ‰