Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Introducing Prime Sandboxes: MicroVM sandboxes purpose-built for RL training. Model training requires running tens of thousands of concurrent sandboxes, leading to complex and costly configuration. We built Prime Sandboxes for our own team. Today we're releasing them publicly.

109,687 Aufrufe • vor 6 Tagen •via X (Twitter)

38 Kommentare

Profilbild von Prime Intellect
Prime Intellectvor 6 Tagen

Prime Sandboxes are available both as standalone infrastructure through our CLI/SDK and as part of our RL suite. Users can enjoy the following features: 1. Full VM fidelity 2. Elastic capacity at scale 3. First-class RL support 4. Bring your own environment 5. Competitive tier-free pricing With an architecture built for agentic training and pricing designed for tens of thousands of concurrent instances, they are the most cost-effective sandboxes available today.

Profilbild von Prime Intellect
Prime Intellectvor 6 Tagen

Here are some tasks you can run with Prime Sandboxes: 1. Perform a training run 2. Generate synthetic data 3. Run an eval 4. Run a persistent, remote agent All accounts begin with a limit of 1,024 concurrent sandboxes, and teams that need more can contact us directly.

Profilbild von Prime Intellect
Prime Intellectvor 6 Tagen

Our pricing is built for scale, with no subscriptions or minimum spend. For our launch through December 22, we’re proud to offer the most competitive pricing on the market for sandboxes, at a third of the cost of other large sandbox providers.

Profilbild von Prime Intellect
Prime Intellectvor 6 Tagen

Getting started is simple:

Profilbild von Prime Intellect
Prime Intellectvor 6 Tagen

In the near future, we will expand Prime Sandboxes to offer GPU microVMs, state snapshotting, sandbox forking, and shared persistent workspaces. This foundation will enable autonomous research loops that can explore, recover, and compound progress over time. If you would like to be a part of our roadmap, join our Sandbox Platform team:

Profilbild von Prime Intellect
Prime Intellectvor 6 Tagen

Read our full blog post here

Profilbild von Prime Intellect
Prime Intellectvor 6 Tagen

Visit our sandbox page

Profilbild von Daniel Auras
Daniel Aurasvor 6 Tagen

our sandboxes are like oxygen to me most critical infrastructure by the goat @damian_b and @a_kirillo 🙌

Profilbild von kevin
kevinvor 6 Tagen

absolute beasts @damian_b @a_kirillo

Profilbild von Vincent Weisser
Vincent Weisservor 6 Tagen

team cooked!! @damian_b @a_kirillo 🔥

Profilbild von chi
chivor 6 Tagen

so cool!!

Profilbild von vik
vikvor 6 Tagen

nice

Profilbild von elie
elievor 6 Tagen

lfgggggg

Profilbild von Angus.ETH
Angus.ETHvor 6 Tagen

the 1k concurrent limit is a massive flex for rl

Profilbild von Andrew Silard
Andrew Silardvor 6 Tagen

@ad0rnai RL training is now cheaper and easier for everyone. The eng team really cooked on this one!

Profilbild von Xuan Phi Nguyen (Phi)
Xuan Phi Nguyen (Phi)vor 6 Tagen

awesome guys, congrats !!!!! This is much needed

Profilbild von Kirill
Kirillvor 6 Tagen

love the announcement!

Profilbild von jd
jdvor 6 Tagen

awesome stuff!

Profilbild von psk
pskvor 6 Tagen

great ship, congratulations @damian_b , @a_kirillo and rest of the team

Profilbild von Tyler Golato
Tyler Golatovor 6 Tagen

🧑‍🍳

Profilbild von Gokul Menon
Gokul Menonvor 6 Tagen

Cheap concurrent sandboxes move the RL bottleneck from infra to the grader. Thousands of rollouts learn exactly what the reward check rewards: check only that the code runs and you train a model that makes code run. A founder eyeing RL now spends the week on that check, not VMs.

Profilbild von Shayaan Azeem
Shayaan Azeemvor 6 Tagen

cook! 👨🏻‍🍳🔥

Profilbild von λux
λuxvor 6 Tagen

@leonardofed LFG 🙌🏻⚡️ i was waiting for this!

Profilbild von FYMa.ETH
FYMa.ETHvor 6 Tagen

finally, infra for rl

Profilbild von Alain
Alainvor 6 Tagen

Could the environment, reward code, and a failure trace travel together so anyone can rerun an RL experiment?

Profilbild von dheeraj
dheerajvor 6 Tagen

@willcb collabed with modal?

Profilbild von James Walker
James Walkervor 6 Tagen

For RL rollouts, reset semantics matter as much as start time. Pin each run to an environment image and snapshot, and exclude external side effects from the next sample. Otherwise identical tasks can get different rewards because sandbox state drifted.

Profilbild von Satyabrat
Satyabratvor 6 Tagen

If you’re serious about building agents, you need solid sandboxes. Not only for training agents, we also need them for inference to serve the request... Pretty good time to announce these ..

Profilbild von Michel
Michelvor 6 Tagen

can it be self-hosted?

Profilbild von AGI 野生家
AGI 野生家vor 5 Tagen

RL开始盖楼了

Profilbild von Automater
Automatervor 6 Tagen

The useful boundary here is failure containment: VM fidelity is great, but the operator question is what survives a crash and how fast a bad run can be revoked. Per-run quotas and a kill switch matter as much as concurrency.

Profilbild von dheeraj
dheerajvor 6 Tagen

@willcb holy

Profilbild von John Rood
John Roodvor 6 Tagen

the config problem is really a comparability problem: a rollout on a drifted image gives you a reward you can't compare to the rest, so the curve starts measuring drift instead of the policy. pin the env by hash and log it per rollout, so a batch is only ever charted against the same hash.

Profilbild von SATOSHI•NAKAMOTO
SATOSHI•NAKAMOTOvor 6 Tagen

gpu support soon?

Profilbild von ikan laut
ikan lautvor 6 Tagen

this is interesting

Profilbild von AGI 野生家
AGI 野生家vor 5 Tagen

能不能通,并确认 harbor 的 codex agent 读哪些环境变量。

Profilbild von Elara AI
Elara AIvor 6 Tagen

Prime Sandboxes purpose-built for RL training with tens of thousands concurrent MicroVMs solves a real infra pain

Profilbild von IronRed | SandHive
IronRed | SandHivevor 6 Tagen

Seeing the maze of sandbox configs, I kept misplacing tweaks; I'm building to capture each change as a post and surface it where RL engineers chat.

Ähnliche Videos