Загрузка видео...

Не удалось загрузить видео

На главную

Introducing Prime Sandboxes: MicroVM sandboxes purpose-built for RL training. Model training requires running tens of thousands of concurrent sandboxes, leading to complex and costly configuration. We built Prime Sandboxes for our own team. Today we're releasing them publicly.

112,320 просмотров • 6 дней назад •via X (Twitter)

Комментарии: 38

Фото профиля Prime Intellect
Prime Intellect6 дней назад

Prime Sandboxes are available both as standalone infrastructure through our CLI/SDK and as part of our RL suite. Users can enjoy the following features: 1. Full VM fidelity 2. Elastic capacity at scale 3. First-class RL support 4. Bring your own environment 5. Competitive tier-free pricing With an architecture built for agentic training and pricing designed for tens of thousands of concurrent instances, they are the most cost-effective sandboxes available today.

Фото профиля Prime Intellect
Prime Intellect6 дней назад

Here are some tasks you can run with Prime Sandboxes: 1. Perform a training run 2. Generate synthetic data 3. Run an eval 4. Run a persistent, remote agent All accounts begin with a limit of 1,024 concurrent sandboxes, and teams that need more can contact us directly.

Фото профиля Prime Intellect
Prime Intellect6 дней назад

Our pricing is built for scale, with no subscriptions or minimum spend. For our launch through December 22, we’re proud to offer the most competitive pricing on the market for sandboxes, at a third of the cost of other large sandbox providers.

Фото профиля Prime Intellect
Prime Intellect6 дней назад

Getting started is simple:

Фото профиля Prime Intellect
Prime Intellect6 дней назад

In the near future, we will expand Prime Sandboxes to offer GPU microVMs, state snapshotting, sandbox forking, and shared persistent workspaces. This foundation will enable autonomous research loops that can explore, recover, and compound progress over time. If you would like to be a part of our roadmap, join our Sandbox Platform team:

Фото профиля Prime Intellect
Prime Intellect6 дней назад

Read our full blog post here

Фото профиля Prime Intellect
Prime Intellect6 дней назад

Visit our sandbox page

Фото профиля Daniel Auras
Daniel Auras6 дней назад

our sandboxes are like oxygen to me most critical infrastructure by the goat @damian_b and @a_kirillo 🙌

Фото профиля kevin
kevin6 дней назад

absolute beasts @damian_b @a_kirillo

Фото профиля Vincent Weisser
Vincent Weisser6 дней назад

team cooked!! @damian_b @a_kirillo 🔥

Фото профиля chi
chi6 дней назад

so cool!!

Фото профиля vik
vik6 дней назад

nice

Фото профиля elie
elie6 дней назад

lfgggggg

Фото профиля Angus.ETH
Angus.ETH6 дней назад

the 1k concurrent limit is a massive flex for rl

Фото профиля Andrew Silard
Andrew Silard6 дней назад

@ad0rnai RL training is now cheaper and easier for everyone. The eng team really cooked on this one!

Фото профиля Xuan Phi Nguyen (Phi)
Xuan Phi Nguyen (Phi)6 дней назад

awesome guys, congrats !!!!! This is much needed

Фото профиля Kirill
Kirill6 дней назад

love the announcement!

Фото профиля jd
jd6 дней назад

awesome stuff!

Фото профиля psk
psk6 дней назад

great ship, congratulations @damian_b , @a_kirillo and rest of the team

Фото профиля Tyler Golato
Tyler Golato6 дней назад

🧑‍🍳

Фото профиля Gokul Menon
Gokul Menon6 дней назад

Cheap concurrent sandboxes move the RL bottleneck from infra to the grader. Thousands of rollouts learn exactly what the reward check rewards: check only that the code runs and you train a model that makes code run. A founder eyeing RL now spends the week on that check, not VMs.

Фото профиля Shayaan Azeem
Shayaan Azeem6 дней назад

cook! 👨🏻‍🍳🔥

Фото профиля λux
λux6 дней назад

@leonardofed LFG 🙌🏻⚡️ i was waiting for this!

Фото профиля FYMa.ETH
FYMa.ETH6 дней назад

finally, infra for rl

Фото профиля Alain
Alain6 дней назад

Could the environment, reward code, and a failure trace travel together so anyone can rerun an RL experiment?

Фото профиля dheeraj
dheeraj6 дней назад

@willcb collabed with modal?

Фото профиля James Walker
James Walker6 дней назад

For RL rollouts, reset semantics matter as much as start time. Pin each run to an environment image and snapshot, and exclude external side effects from the next sample. Otherwise identical tasks can get different rewards because sandbox state drifted.

Фото профиля Satyabrat
Satyabrat6 дней назад

If you’re serious about building agents, you need solid sandboxes. Not only for training agents, we also need them for inference to serve the request... Pretty good time to announce these ..

Фото профиля Michel
Michel6 дней назад

can it be self-hosted?

Фото профиля AGI 野生家
AGI 野生家6 дней назад

RL开始盖楼了

Фото профиля Automater
Automater6 дней назад

The useful boundary here is failure containment: VM fidelity is great, but the operator question is what survives a crash and how fast a bad run can be revoked. Per-run quotas and a kill switch matter as much as concurrency.

Фото профиля dheeraj
dheeraj6 дней назад

@willcb holy

Фото профиля John Rood
John Rood6 дней назад

the config problem is really a comparability problem: a rollout on a drifted image gives you a reward you can't compare to the rest, so the curve starts measuring drift instead of the policy. pin the env by hash and log it per rollout, so a batch is only ever charted against the same hash.

Фото профиля SATOSHI•NAKAMOTO
SATOSHI•NAKAMOTO6 дней назад

gpu support soon?

Фото профиля ikan laut
ikan laut6 дней назад

this is interesting

Фото профиля AGI 野生家
AGI 野生家6 дней назад

能不能通,并确认 harbor 的 codex agent 读哪些环境变量。

Фото профиля Elara AI
Elara AI6 дней назад

Prime Sandboxes purpose-built for RL training with tens of thousands concurrent MicroVMs solves a real infra pain

Фото профиля IronRed | SandHive
IronRed | SandHive6 дней назад

Seeing the maze of sandbox configs, I kept misplacing tweaks; I'm building to capture each change as a post and surface it where RL engineers chat.

Похожие видео