
Aesther
@AestherML • 62,464 subscribers
AI Researcher • Accelerationist
Shorts
Okay, I am fairly confident in my hypothesis now. The secret sauce behind Qwen 3.8 27B becomes almost immediately evident during testing. It is not the training data. In fact, I doubt any SFT was involved at all. The model was simply allowed GRPO with a more liberal reasoning context length. As I have explained to you before, distillation through GRPO is powerful, and in certain contexts vastly superior to K/L divergence distillation. Qwen used the outputs of Qwen 3.8 Max as the reinforcement learning objective for Qwen 3.8 27B and told the model reason as much as it has to in order to match the output. Remember, Intelligence is solution, not compression. This is like nesting the output of a function as an additional input to itself, and it is absolutely brilliant.
122,746 Aufrufe