正在加载视频...

视频加载失败

1/ Introducing RL Swarm’s new backend: GenRL. A modular reinforcement learning library built for distributed, fault-tolerant training - now powering RL Swarm from the ground up. 🧵

82,191 次观看 • 1 年前 •via X (Twitter)

10 条评论

gensyn 的头像
gensyn1 年前

2/ Each worker runs its own environment instance, contributes asynchronously to a shared rollout buffer, and updates its model weights independently, so no central controller is required.

gensyn 的头像
gensyn1 年前

3/ GenRL allows RL Swarm to work with any environment, described intuitively through code. This launch incorporates Reasoning Gym out-of-the-box, giving access to >100 community-created environments with no extra configuration required.

gensyn 的头像
gensyn1 年前

4/ What’s new: – Modular GenRL backend – Expanded configuration surface – Prebuilt Docker image for easy deployment – Reasoning Gym environment to enhance model reasoning capabilities – New multi-task swarm

gensyn 的头像
gensyn1 年前

5/ Now live on the Gensyn testnet. You can run RL-Swarm with GenRL today. Full code + setup:

gensyn 的头像
gensyn1 年前

6/ A node update is required for GenRL. Please visit ⁠support-discussion in the Discord if you have any questions.

Gautamgg 🕵 的头像
Gautamgg 🕵1 年前

I want to ask 1 que What about previous trained model data rewards & participants bec it's not showing Is that data saved in your database? @fenbielding @_jamico @_grieve waiting for ans 💙

Mintair | One Click Node🪄 的头像
Mintair | One Click Node🪄1 年前

Looks really interesting, we gotta setup our own custom environment.

AJDominic (🐱,🐐) 的头像
AJDominic (🐱,🐐)1 年前

What gensyn cooking is unmatched!

Bitduke 的头像
Bitduke1 年前

Cool, cool - more modularity

lior.eth (Lior Messika) 的头像
lior.eth (Lior Messika)1 年前

These retro vibes are everything I ever wanted from an AI lab

相关视频

🚨 RL for LLMs is finally accessible. Introducing OpenTinker: The first community-driven, open-source framework designed to democratize Reinforcement Learning for LLMs. Inspired by Thinking Machines's amazing Tinker, we realize the biggest bottleneck in agentic LLM research isn’t the math—it’s the setup. Current RL pipelines are messy. Configuring VeRL for every single experiment is a productivity killer. OpenTinker fixed it. 🛠 How OpenTinker Works: Decoupled Design of Server and Client - Setup Once, Run Forever: Configure the OpenTinker backend on your GPU cluster once. - Develop Locally: Define your RL environments directly on your laptop. - Train on the Cloud: Simply point your local client to the backend. The cluster handles the compute; you handle the science. 📉 The 10x Development Efficiency Thanks to our elegant architectural decomposition, OpenTinker reduces the time to develop a new RL training pipeline by at least an order of magnitude. ⚡ Turn Idle GPU Compute into Gold Small labs often have underutilized hardware. OpenTinker turns your idle GPUs into an internal/external API service for - RL Training - SFT - Inference 🎯 Who needs OpenTinker? - Researchers tired of infrastructure hell. - Labs needing to standardize workflows. - Teams wanting to maximize hardware ROI. Thanks my amazing PhD student Siqi Zhu for leading the project. We are building the future of open RL infra. Be the first to build with us. 👇 Start Building with OpenTinker Now 🚀 Repo: 🌐 Blog: If you believe RL should be accessible to everyone, give us a star, repost this 🔄 post, and let us know what agents you plan to build!

Jiaxuan You

58,258 次观看 • 8 个月前