Astraia's banner
Astraia's profile picture

Astraia

@AstraiaAI61,276 subscribers

AI Researcher & Theorist | Biological and Artificial Neural Networks • Accelerationist • Pro Open-Weights !! 🤗

Shorts

Okay, I am fairly confident in my hypothesis now. The secret sauce behind Qwen 3.8 27B becomes almost immediately evident during testing. It is not the training data. In fact, I doubt any SFT was involved at all. The model was simply allowed GRPO with a more liberal reasoning context length. As I have explained to you before, distillation through GRPO is powerful, and in certain contexts vastly superior to K/L divergence distillation. Qwen used the outputs of Qwen 3.8 Max as the reinforcement learning objective for Qwen 3.8 27B and told the model reason as much as it has to in order to match the output. Remember, Intelligence is solution, not compression. This is like nesting the output of a function as an additional input to itself, and it is absolutely brilliant.

Okay, I am fairly confident in my hypothesis now. The secret sauce behind Qwen 3.8 27B becomes almost immediately evident during testing. It is not the training data. In fact, I doubt any SFT was involved at all. The model was simply allowed GRPO with a more liberal reasoning context length. As I have explained to you before, distillation through GRPO is powerful, and in certain contexts vastly superior to K/L divergence distillation. Qwen used the outputs of Qwen 3.8 Max as the reinforcement learning objective for Qwen 3.8 27B and told the model reason as much as it has to in order to match the output. Remember, Intelligence is solution, not compression. This is like nesting the output of a function as an additional input to itself, and it is absolutely brilliant.

121,447 次观看