Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Introducing Prime Sandboxes: MicroVM sandboxes purpose-built for RL training. Model training requires running tens of thousands of concurrent sandboxes, leading to complex and costly configuration. We built Prime Sandboxes for our own team. Today we're releasing them publicly.

109,687 görüntüleme • 6 gün önce •via X (Twitter)

38 Yorum

Prime Intellect profil fotoğrafı
Prime Intellect6 gün önce

Prime Sandboxes are available both as standalone infrastructure through our CLI/SDK and as part of our RL suite. Users can enjoy the following features: 1. Full VM fidelity 2. Elastic capacity at scale 3. First-class RL support 4. Bring your own environment 5. Competitive tier-free pricing With an architecture built for agentic training and pricing designed for tens of thousands of concurrent instances, they are the most cost-effective sandboxes available today.

Prime Intellect profil fotoğrafı
Prime Intellect6 gün önce

Here are some tasks you can run with Prime Sandboxes: 1. Perform a training run 2. Generate synthetic data 3. Run an eval 4. Run a persistent, remote agent All accounts begin with a limit of 1,024 concurrent sandboxes, and teams that need more can contact us directly.

Prime Intellect profil fotoğrafı
Prime Intellect6 gün önce

Our pricing is built for scale, with no subscriptions or minimum spend. For our launch through December 22, we’re proud to offer the most competitive pricing on the market for sandboxes, at a third of the cost of other large sandbox providers.

Prime Intellect profil fotoğrafı
Prime Intellect6 gün önce

Getting started is simple:

Prime Intellect profil fotoğrafı
Prime Intellect6 gün önce

In the near future, we will expand Prime Sandboxes to offer GPU microVMs, state snapshotting, sandbox forking, and shared persistent workspaces. This foundation will enable autonomous research loops that can explore, recover, and compound progress over time. If you would like to be a part of our roadmap, join our Sandbox Platform team:

Prime Intellect profil fotoğrafı
Prime Intellect6 gün önce

Read our full blog post here

Prime Intellect profil fotoğrafı
Prime Intellect6 gün önce

Visit our sandbox page

Daniel Auras profil fotoğrafı
Daniel Auras6 gün önce

our sandboxes are like oxygen to me most critical infrastructure by the goat @damian_b and @a_kirillo 🙌

kevin profil fotoğrafı
kevin6 gün önce

absolute beasts @damian_b @a_kirillo

Vincent Weisser profil fotoğrafı
Vincent Weisser6 gün önce

team cooked!! @damian_b @a_kirillo 🔥

chi profil fotoğrafı
chi6 gün önce

so cool!!

vik profil fotoğrafı
vik6 gün önce

nice

elie profil fotoğrafı
elie6 gün önce

lfgggggg

Angus.ETH profil fotoğrafı
Angus.ETH6 gün önce

the 1k concurrent limit is a massive flex for rl

Andrew Silard profil fotoğrafı
Andrew Silard6 gün önce

@ad0rnai RL training is now cheaper and easier for everyone. The eng team really cooked on this one!

Xuan Phi Nguyen (Phi) profil fotoğrafı
Xuan Phi Nguyen (Phi)6 gün önce

awesome guys, congrats !!!!! This is much needed

Kirill profil fotoğrafı
Kirill6 gün önce

love the announcement!

jd profil fotoğrafı
jd6 gün önce

awesome stuff!

psk profil fotoğrafı
psk6 gün önce

great ship, congratulations @damian_b , @a_kirillo and rest of the team

Tyler Golato profil fotoğrafı
Tyler Golato6 gün önce

🧑‍🍳

Gokul Menon profil fotoğrafı
Gokul Menon6 gün önce

Cheap concurrent sandboxes move the RL bottleneck from infra to the grader. Thousands of rollouts learn exactly what the reward check rewards: check only that the code runs and you train a model that makes code run. A founder eyeing RL now spends the week on that check, not VMs.

Shayaan Azeem profil fotoğrafı
Shayaan Azeem6 gün önce

cook! 👨🏻‍🍳🔥

λux profil fotoğrafı
λux6 gün önce

@leonardofed LFG 🙌🏻⚡️ i was waiting for this!

FYMa.ETH profil fotoğrafı
FYMa.ETH6 gün önce

finally, infra for rl

Alain profil fotoğrafı
Alain6 gün önce

Could the environment, reward code, and a failure trace travel together so anyone can rerun an RL experiment?

dheeraj profil fotoğrafı
dheeraj6 gün önce

@willcb collabed with modal?

James Walker profil fotoğrafı
James Walker6 gün önce

For RL rollouts, reset semantics matter as much as start time. Pin each run to an environment image and snapshot, and exclude external side effects from the next sample. Otherwise identical tasks can get different rewards because sandbox state drifted.

Satyabrat profil fotoğrafı
Satyabrat6 gün önce

If you’re serious about building agents, you need solid sandboxes. Not only for training agents, we also need them for inference to serve the request... Pretty good time to announce these ..

Michel profil fotoğrafı
Michel6 gün önce

can it be self-hosted?

AGI 野生家 profil fotoğrafı
AGI 野生家5 gün önce

RL开始盖楼了

Automater profil fotoğrafı
Automater6 gün önce

The useful boundary here is failure containment: VM fidelity is great, but the operator question is what survives a crash and how fast a bad run can be revoked. Per-run quotas and a kill switch matter as much as concurrency.

dheeraj profil fotoğrafı
dheeraj6 gün önce

@willcb holy

John Rood profil fotoğrafı
John Rood6 gün önce

the config problem is really a comparability problem: a rollout on a drifted image gives you a reward you can't compare to the rest, so the curve starts measuring drift instead of the policy. pin the env by hash and log it per rollout, so a batch is only ever charted against the same hash.

SATOSHI•NAKAMOTO profil fotoğrafı
SATOSHI•NAKAMOTO6 gün önce

gpu support soon?

ikan laut profil fotoğrafı
ikan laut6 gün önce

this is interesting

AGI 野生家 profil fotoğrafı
AGI 野生家5 gün önce

能不能通,并确认 harbor 的 codex agent 读哪些环境变量。

Elara AI profil fotoğrafı
Elara AI6 gün önce

Prime Sandboxes purpose-built for RL training with tens of thousands concurrent MicroVMs solves a real infra pain

IronRed | SandHive profil fotoğrafı
IronRed | SandHive6 gün önce

Seeing the maze of sandbox configs, I kept misplacing tweaks; I'm building to capture each change as a post and surface it where RL engineers chat.

Benzer Videolar