Aesther's banner
Aesther's profile picture

Aesther

@AestherML62,287 subscribers

AI Researcher & Theorist | Biological and Artificial Neural Networks • Accelerationist 🤗🧸

Shorts

Okay, I am fairly confident in my hypothesis now. The secret sauce behind Qwen 3.8 27B becomes almost immediately evident during testing. It is not the training data. In fact, I doubt any SFT was involved at all. The model was simply allowed GRPO with a more liberal reasoning context length. As I have explained to you before, distillation through GRPO is powerful, and in certain contexts vastly superior to K/L divergence distillation. Qwen used the outputs of Qwen 3.8 Max as the reinforcement learning objective for Qwen 3.8 27B and told the model reason as much as it has to in order to match the output. Remember, Intelligence is solution, not compression. This is like nesting the output of a function as an additional input to itself, and it is absolutely brilliant.

Okay, I am fairly confident in my hypothesis now. The secret sauce behind Qwen 3.8 27B becomes almost immediately evident during testing. It is not the training data. In fact, I doubt any SFT was involved at all. The model was simply allowed GRPO with a more liberal reasoning context length. As I have explained to you before, distillation through GRPO is powerful, and in certain contexts vastly superior to K/L divergence distillation. Qwen used the outputs of Qwen 3.8 Max as the reinforcement learning objective for Qwen 3.8 27B and told the model reason as much as it has to in order to match the output. Remember, Intelligence is solution, not compression. This is like nesting the output of a function as an additional input to itself, and it is absolutely brilliant.

122,423 Aufrufe