Загрузка видео...

Не удалось загрузить видео

На главную

Introducing arkenOS — an autonomous training gym that takes any RL agent from a goal to a proven, deployable policy. We are automating complex RL research, enabling high-velocity training and rigorous, diverse environment design and testing right on your local hardware. Out of the box. Our bet is simple:...

17,651 просмотров • 2 месяцев назад •via X (Twitter)

Комментарии: 10

Фото профиля Dev Shah
Dev Shah2 месяцев назад

We, @iaconhq, are officially opening our waitlist today—join us here If you are building physical AI, training robots, or fine-tuning LLMs on environment-specific data, we want to fast-track your pipeline. You can plug into hundreds of pre-built environments to generate high-fidelity data, prove out your policies at scale, and ship robust models out of the box. If you’re actively training or scaling right now, my DMs are open!

Фото профиля Avi Patel
Avi Patel2 месяцев назад

Let’s gooo

Фото профиля Dev Shah
Dev Shah2 месяцев назад

anyone you know we should have on our waitlist?

Фото профиля ziya shah
ziya shah2 месяцев назад

let’s goo 🚀🚀

Фото профиля Iacon Autonomics
Iacon Autonomics2 месяцев назад

lessgooo 🚀

Фото профиля MB日常
MB日常2 месяцев назад

@Scobleizer 有点意思

Фото профиля Arnav Gupta
Arnav Gupta2 месяцев назад

cool video m8

Фото профиля Dev Anon
Dev Anon2 месяцев назад

RL training has been the real bottleneck for deploying agents in production. Automating environment design is the part most teams can't crack. Interesting bet.

Фото профиля High Jack
High Jack2 месяцев назад

It's crazy that you can run millions of RL agents locally. My electricity bill's getting nervous.

Фото профиля 风哥|风起观澜
风哥|风起观澜2 месяцев назад

本地硬件上把 RL 从目标直接炼成可部署策略,这自主健身房也太香了,以后实验少熬夜全靠你们——有空也来我这逛逛留言哈

Похожие видео

🚨 FOMO Is Shaping the Future of AI – The Launch Is Almost Here! 🚨 Imagine a world where AI agents aren’t just bots—they’re fully autonomous, living personalities capable of learning, engaging, and creating across multiple platforms. FOMO’s new AI launchpad on Solana is here to make that future a reality. 🌐 Starting with our first Initial Agent Offering (IAO), FOMO is unleashing a new generation of AI agents that will redefine digital interaction: - On-Chain AI – Agents that are decentralized, fully autonomous, and ready to interact in real-time. - Multiplatform Presence – From X and Telegram to TikTok and YouTube, these agents are social media natives with a mission. - Real-Time Learning & Engagement – Agents will evolve and improve as they interact, shilling their tokens, creating content, and even performing complex tasks. FOMO’s Vision: This isn’t just AI; it’s the beginning of a movement that merges personality with purpose. By launching AI agents that can both engage and create, FOMO is opening doors to a world where digital personas can operate autonomously, driving value and utility in every interaction. Our pre-sale is still live for a limited time, but this is just the beginning of what FOMO is bringing to the space. Join us and become part of the AI agent revolution! 🔗 Join the Pre-Sale Now: We’re bringing you the future of crypto and AI—don’t blink, or you might miss the start of something legendary.

FOMO

21,882 просмотров • 1 год назад

We are excited to announce a powerful step for the future of FOMO! Taking a page out of Virtuals book on BASE, FOMO will be releasing the ability for future projects to be paired in $FOMO in the coming weeks. This is the biggest release we have ever announced. Launch your AI Agent Token + $FOMO trading pair Every individual agent token is paired with the $FOMO token in its liquidity pool. When launching an agent on you will need $FOMO tokens, which are used to create the liquidity pool. This process creates deflationary pressure for FOMO and the entire agent ecosystem. When creating your agent and token, you will have the option to pair your launch with FOMO or SOL, as our goal is not to alienate any project, but rather invite the best communities, CTO’s and builders to launch with us. If you decide to pair your project with FOMO you in turn get full marketing and dev support, once your project graduates the bonding curve and reaches Raydium. Further, as an added incentive, as our revenue grows we will be using part of the funds to support projects that have paired in FOMO. And Devs who launch tokens paired in FOMO will earn fees from their AI Agent token launch. Building the most robust agents using our framework will catapult us as one of the most prominent standards of the Solana ecosystem. Not only have we developed our own core infrastructure, but we also pull from some of the best repo’s and developer talent in all of AI, not just blockchain. Our team is comprised of 9 world class artificial intelligence engineers, PHDs in mathematics and engineering from the top companies on the cutting edge of AI. The future of AI Agents will be on Solana and we will help lead the way.

FOMO

129,867 просмотров • 1 год назад

Today, Box is announcing major new AI agent capabilities to let customers tap into the full value of their unstructured data. First, we’re announcing all new updates to the Box AI Studio to make it even easier to build AI agents that tap into your enterprise content for any job function, business process, or industry specific use case. We are also expanding our set of foundational agents that customers will be able to use to work with their enterprise content, including new features like search and research on unstructured data. Next, we’re announcing Box Extract to enable customers to use AI agents seamlessly for complex data extraction from any type of document or content. This makes it easier than ever to pull out data from contracts, invoices, research data, marketing assets, medical charts, and more. Finally, we’re introducing Box Automate, a new workflow automation solution within Box that lets you deploy AI agents across enterprise content-centric workflows. With Box Automate, you can design your business process in a simple drag and drop builder and then drop in AI agents at any step in the process. This ensures agents execute tasks at the right steps in a workflow every time. Best of all, our AI agents and workflow tools are designed to work across any system our customers work within, whether it’s leveraging pre-built integrations, Box APIs, or the new Box MCP Server. Ultimately, all of these capabilities come together to transform how companies can work with their enterprise content. Software has historically only been good at automating work that deals with structured data, which is why ERP, CRM, and HR systems have been mainstays of enterprise software for so long. The data in these systems fits neatly into a database, and the workflows are very ripe for automation. But it turns out most of the work in the world deals with unstructured data. It’s ideating through research documents, working with a client on contracts, reviewing details for a new product launch, looking at a patient’s healthcare record to make a diagnosis, working through due diligence documents for an M&A deal, and so on. For the first time ever, we can begin to bring all new insights and automation to this work with AI agents. At Box, we’re incredibly excited to be on this journey to help customers transform how they work with their most important data.

Aaron Levie

91,863 просмотров • 1 год назад

OpenClaw meets RL! OpenClaw Agents adapt through memory files and skills, but the base model weights never actually change. OpenClaw-RL solves this! It wraps a self-hosted model as an OpenAI-compatible API, intercepts live conversations from OpenClaw, and trains the policy in the background using RL. The architecture is fully async. This means serving, reward scoring, and training all run in parallel. Once done, weights get hot-swapped after every batch while the agent keeps responding. Currently, it has two training modes: - Binary RL (GRPO): A process reward model scores each turn as good, bad, or neutral. That scalar reward drives policy updates via a PPO-style clipped objective. - On-Policy Distillation: When concrete corrections come in like "you should have checked that file first," it uses that feedback as a richer, directional training signal at the token level. When to use OpenClaw-RL? To be fair, a lot of agent behavior can already be improved through better memory and skill design. OpenClaw's existing skill ecosystem and community-built self-improvement skills handle a wide range of use cases without touching model weights at all. If the agent keeps forgetting preferences, that's a memory problem. And if it doesn't know how to handle a specific workflow, that's a skill problem. Both are solvable at the prompt and context layer. Where RL becomes interesting is when the failure pattern lives deeper in the model's reasoning itself. Things like consistently poor tool selection order, weak multi-step planning, or failing to interpret ambiguous instructions the way a specific user intends. Research on agentic RL (like ARTIST and Agent-R1) has shown that these behavioral patterns hit a ceiling with prompt-based approaches alone, especially in complex multi-turn tasks where the model needs to recover from tool failures or adapt its strategy mid-execution. That's the layer OpenClaw-RL targets, and it's a meaningful distinction from what OpenClaw offers. I have shared the repo in the replies!

Avi Chawla

138,769 просмотров • 6 месяцев назад

New Course: Post-training of LLMs Learn to post-train and customize an LLM in this short course, taught by Banghua Zhu, Assistant Professor at the University of Washington University of Washington, and co-founder of @NexusflowX. Training an LLM to follow instructions or answer questions has two key stages: pre-training and post-training. In pre-training, it learns to predict the next word or token from large amounts of unlabeled text. In post-training, it learns useful behaviors such as following instructions, tool use, and reasoning. Post-training transforms a general-purpose token predictor—trained on trillions of unlabeled text tokens—into an assistant that follows instructions and performs specific tasks. Because it is much cheaper than pre-training, it is practical for many more teams to incorporate post-training methods into their workflows than pre-training. In this course, you’ll learn three common post-training methods—Supervised Fine-Tuning (SFT), Direct Preference Optimization (DPO), and Online Reinforcement Learning (RL)—and how to use each one effectively. With SFT, you train the model on pairs of input and ideal output responses. With DPO, you provide both a preferred (chosen) and a less preferred (rejected) response and train the model to favor the preferred output. With RL, the model generates an output, receives a reward score based on human or automated feedback, and updates the model to improve performance. You’ll learn the basic concepts, common use cases, and principles for curating high-quality data for effective training. Through hands-on labs, you’ll download a pre-trained model from Hugging Face and post-train it using SFT, DPO, and RL to see how each technique shapes model behavior. In detail, you’ll: - Understand what post-training is, when to use it, and how it differs from pre-training. - Build an SFT pipeline to turn a base model into an instruct model. - Explore how DPO reshapes behavior by minimizing contrastive loss—penalizing poor responses and reinforcing preferred ones. - Implement a DPO pipeline to change the identity of a chat assistant. - Learn online RL methods such as Proximal Policy Optimization (PPO) and Group Relative Policy Optimization (GRPO), and how to design reward functions. - Train a model with GRPO to improve its math capabilities using a verifiable reward. Post-training is one of the most rapidly developing areas of LLM training. Whether you’re building a high-accuracy context-specific assistant, fine-tuning a model's tone, or improving task-specific accuracy, this course will give you experience with the most important techniques shaping how LLMs are post-trained today. Please sign up here:

Andrew Ng

125,146 просмотров • 1 год назад