Loading video...
Video Failed to Load
New high quality blog post: "How to train a Frontier-Level world model" in collaboration with reactor This is going to be very useful for practitioners
97,514 views • 2 months ago •via X (Twitter)
40 Comments

We trained a 1.6B-parameter Minecraft world model and today we’re open-sourcing the model and code. We wrote the guide we wish we’d had: architecture, scaling, and the stability of the algorithm You can play it live:

@reactorworld Congrats and your diagrams are elite

@reactorworld Thanks man 💪

@reactorworld spectacular!!!

@reactorworld Thanks!

@reactorworld This is so cool, thank you for publishing this!

@reactorworld Thanks!

@reactorworld Super cool Francesco!

@reactorworld Thanks!

@reactorworld The dedication it takes to do this is incredible. Upmost respect man!

@reactorworld ❤️

@reactorworld Amazing work

@reactorworld This is already on my TO-READ list for next flight! Great job also on keeping the article brief. This certainly wasn't easy.

@reactorworld Super cool! Thank you so much :)

@reactorworld great work!

@reactorworld thanks!

@reactorworld it's been amazing working with you on this @FrancescoSacco1 !!!!

@reactorworld Pleasure is mine! thanks a lot guys!

@reactorworld Will read this asap

@reactorworld Thank you for opening this up. Minecraft for education is a powerful tool.

@reactorworld Oh, that’s nice. I can’t wait to see more and more open-source/weight world models and beyond

@reactorworld This is insane thanks for your hardwork

@reactorworld really sick as always!! are you considering also writing about any ablations you may have done?

@reactorworld Great detailed writeup!

@reactorworld thanks!

@reactorworld Very interesting implementation of Dreamer 4! I notice that the model seems to struggle with remembering objects out of view. If I see a body of water and look away for 1s, the water turns into land. Do you think it's due to the short context window or the model's limitation?

@reactorworld The context windows is of 192 frames, which should be about 9 seconds So the problem is clearly model capacity. The thing is that MSE-loss is probably not the right way to teach the model to remember stuff. There needs to be a different loss for that imo

@reactorworld this is really cool

I am also working on training a world model and trying to focus really on consistency. I have few notes to share: 1. I found its essential to keep rendering and world model completely separate. Combining their loss really really derails the world model while the loss goes down (as you mentioned also). 2. I tried all the autoencoders but at the end of the day, simplest state representation which can be used to recreate a frame works best. I am not even talking about autoencoder, but a coded compositor. 3. Make the state representation as small as possible! Upon much analysis I found that many times even redundant information in the state is enough to cause issues. Atleast from my experience, I can not stress enough that state representation should be as small as humanly possible. 4. Data! I see that you also used rl agent to generate data, but that is actually not ideal. The rl agent essentially tries to reach nash equilibrium. But to create a world model we want balanced data about all the possible events that can happen. rl agents don't help with that. More often than not, they will play the optimal moves and ignore an entire set of world mechanics. 5. Frames look good but video clips are kinda shit. Most difficult problem of this entire exercise. The problem with self rollout is that the model is trained on behaviour cloning, but since you are doing self rollout, the model's output will accumulate error over frames. So a model with excellent loss will also perform poorly on self rollout. You really really need to train the model on noisy data, not perfect behaviour cloning. 6. Closed loop validation is the only honest metric. Everything else is anti signal.

1. You mean that you tried training the dynamics model and tokenizer at the same time? 2. how do you explain the results from this paper here? . what were your results? 3. I guess it depends, there is this paper that got some really interesting results . does it check out with yours? 4. Actually we used minecraft vpt dataset. However we still had to remove some pieces of bad data (such as the player standing still for days) 5. We didn't do behaviour cloning for minecraft, unfortunately we only had time to do it in coinrun :( 6. 100% agree there are no universally good evals for world models

1. Actually many combinations one of which was training them together. Essentially the input to the model are image patches , and the output is the next frame using flow matching. Other combos were pretrained autoencoder, fine tuned autoencoder, from scratch trained autoencoder. But all of them had artifacts in the image quality. None of them can give 100% accurate frames, which causes issues during rollout (as in next generated frame becomes input in next loop) 2. I did try something like that. It actually had best results out of all the other techniques. But when its compared head-to-head against a factorized sprite/layer compositor, the factorized compositor beat it decisively. MAETok is perf celing when vision is the only signal, but if you can do it mechanically, as in maybe you have player position, and the minecraft seed, you can mechanically recreate the frame 1 to 1 with absolutely 0 error. It does feel somewhat like cheating bcs now rendering is not part of the neural network, but this makes the signal to the world model essentially error free and allows much longer rollout using the world model. The world model would still have to learn gravity and and stuff, but the rendering part is error free. 3. pretty much one to one

@reactorworld That’s amazing

@reactorworld What was the biggest obstacle to getting it to work?

@reactorworld legend!

Read this properly this morning, genuinely good writeup. I've been building the full V4 loop in PyTorch at small scale (tokenizer → dynamics → agent finetune → imagination RL). A few things that bit me hard in the phases you have on the roadmap, in case they save you time: • Phase 2: Initially I detached the agent embedding from the world model. I had a .detach() there and the policy was silently action-blind every loss curve looked healthy, action conditioning was +0.5%. • Bootstrap loss raw values of 1000–9000 at K_max=64 are expected, not a bug. The (1−τ−d/2) divisor blows up near the boundary. Repo if it's useful: — different scale and domain (dm_control, ~47M params), so a lot won't transfer to Minecraft directly.

@reactorworld Where is it?

@reactorworld

@reactorworld bookmarking this to pretend i understand how world models actually work 🤫

@reactorworld Hoping there's an eval section. That's where most world model writeups suddenly go quiet.

@reactorworld You can literally play in real time with the model. That's the real eval. Most evals in WMs are flawed in one way or another.

