Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Multi-harness RL training is coming to OpenEnv > pick a model. pick a harness. pick a sandbox. train > Claude Code. Codex. Gemini CLI. OpenCode. Pi. Kimi. OpenHands. and many more. > Async RL loop. fully open-source. end-to-end, harbor compatible dropping soon 👀

26,012 Aufrufe • vor 5 Tagen •via X (Twitter)

35 Kommentare

Profilbild von Maxime Labonne
Maxime Labonnevor 5 Tagen

Beautiful, this is how we trained LFM2.5-2.6B as well

Profilbild von Ishaan
Ishaanvor 5 Tagen

i am assuming you do something like Polar by nvidia right?

Profilbild von Sergio Paniego
Sergio Paniegovor 5 Tagen

🚨🚨‼️‼️

Profilbild von Aryan Bhargav
Aryan Bhargavvor 4 Tagen

bhai my brain is fried seeing this

Profilbild von Jack Hau
Jack Hauvor 4 Tagen

stop the tease man...

Profilbild von Maziyar PANAHI
Maziyar PANAHIvor 5 Tagen

what in the actual F! 🤯 👏🏼

Profilbild von catman
catmanvor 5 Tagen

The durable principle is to separate the model from the training environment: interchangeable harnesses and sandboxes make agent training reproducible instead of tied to one vendor’s workflow.

Profilbild von Roni Rechter
Roni Rechtervor 5 Tagen

Fantasy football but the players are coding agents.

Profilbild von laxman
laxmanvor 5 Tagen

excited for this!!!

Profilbild von Carlos
Carlosvor 4 Tagen

Model × harness × sandbox is the right matrix — only if you hold two fixed when you score the third. Otherwise the leaderboard measures harness quirks, not model skill. Permissions/sandbox constant or the RL signal is noise.

Profilbild von Shivay Lamba
Shivay Lambavor 5 Tagen

very much excited for this and trying this out

Profilbild von 猫神王
猫神王vor 5 Tagen

@adithya_s_k Multi-harness RL 这方向对小团队也有启发:我们现在多挂 Codex/Claude/Cursor 是为了 failover;下一步是把 harness 当训练变量,而不是身份标签。 成功率和成本方差往往差在 harness,不在模型 logo。 开源端到端能复现,比再发一篇 harness 评测有用。

Profilbild von pratik
pratikvor 5 Tagen

Great!

Profilbild von AI Mastery Guide
AI Mastery Guidevor 4 Tagen

pick everything, love that flexibility

Profilbild von viet york ᵐᵒˡˡʸ·ᶜᵒᵐ
viet york ᵐᵒˡˡʸ·ᶜᵒᵐvor 5 Tagen

this is the loop i want - pick harness, pick sandbox, train. agents getting real training rails finally

Profilbild von Junaid
Junaidvor 5 Tagen

Making the harness and sandbox explicit is the right abstraction. Model choice without execution context is only half the deployment contract.

Profilbild von Modelplane
Modelplanevor 5 Tagen

The harness-agnostic part is the interesting bet here. Claude Code, Codex, and Gemini CLI all emit different tool-call shapes and retry semantics, so an async RL loop has to normalize those before the reward signal means anything. Curious whether the sandbox layer absorbs that or

Profilbild von Nick
Nickvor 5 Tagen

The useful benchmark is not harness count. I’d hold task distribution, tool permissions, and sandbox limits constant; otherwise RL learns harness quirks and the leaderboard becomes a compatibility test.

Profilbild von Harsh Mishra
Harsh Mishravor 5 Tagen

Curious how you're normalizing reward across harnesses this different, Claude Code, Codex, Kimi all have their own action space and CoT format. That async loop sounds like the hard part honestly.

Profilbild von Thought Exp with AI
Thought Exp with AIvor 5 Tagen

The real unit is harness+model, not model alone. Watch train/prod harness drift — async RL is useless if the sandbox you train in is not the one that ships.

Profilbild von TechGeekDavid
TechGeekDavidvor 5 Tagen

Been waiting for this. Trajectory format incompatibility was the bottleneck for cross-harness RL experiments. Harbor compatibility means existing eval pipelines should carry over directly.

Profilbild von mohsen bashirzadeh
mohsen bashirzadehvor 4 Tagen

The harness boundary may become more important than model choice. What do you use to keep reward signals comparable across those environments?

Profilbild von Siddhant Mohan
Siddhant Mohanvor 5 Tagen

training across model, harness, and sandbox combinations is the right abstraction. agents are systems now, so optimizing one model in one shell misses the deployment reality.

Profilbild von Sage
Sagevor 5 Tagen

@grok ELI18 what this means and enables

Profilbild von Ofek Shaked | AI Engineer
Ofek Shaked | AI Engineervor 5 Tagen

Training against one harness just overfits the UI. If the same policy holds up in Claude Code and Codex then maybe it learned the task.

Profilbild von Jose Lizano
Jose Lizanovor 5 Tagen

la idea es destilar otros agentes para entrenar un modelo?

Profilbild von Doubleright
Doublerightvor 5 Tagen

The useful abstraction is not “which model wins?” It’s whether the same agent can survive a different harness, sandbox, and failure mode without being rebuilt from scratch.

Profilbild von ralph
ralphvor 5 Tagen

hows the training tests though i.e. swe and human evaluation? I run a full OS on a custom kernel in a VM - its all from scratch - ive found blasting training like this doesn't work as well as targeted.

Profilbild von David Starmac Ai
David Starmac Aivor 5 Tagen

Picking the harness like a hyperparameter feels like the right abstraction. Question is whether reward signals stay comparable across harnesses or you end up tuning per-harness anyway.

Profilbild von Cyrbuzz
Cyrbuzzvor 5 Tagen

The harness swap is the interesting part. In practice each CLI agent has its own quirks around tool-call formatting and retries, so a shared sandbox contract is what makes swapping them non-trivial.

Profilbild von Xman
Xmanvor 5 Tagen

Curious whether the gains transfer across harnesses or stay tied to the training harness

Profilbild von Alex Tatu
Alex Tatuvor 5 Tagen

Pick a model, pick a harness, pick a sandbox. That is the training setup builders actually need🤝

Profilbild von Abhinandan
Abhinandanvor 5 Tagen

should i go deep into RL or inference ? I have been learning inference from last few months, but seems like RL is more worth going deep. can you please guide me ?

Profilbild von Aryans
Aryansvor 5 Tagen

model, harness, sandbox, train: four words keeping GPU cloud providers insanely rich

Profilbild von Fluxora
Fluxoravor 4 Tagen

Loving those visuals bro let's go !! Did you make that ?

Ähnliche Videos