正在加载视频...
视频加载失败
still experimenting with LoRA based on the Thinking Machines configuration and just implemented it in colab. In this notebook I set up a fine tune of Qwen/Qwen3-0.6B on the OpenR1-Math dataset with lora rank of 1. with this setup you can get the same reward accuracy as full fine-tuning,... show more
0 条评论
暂无评论
原始帖子的评论将显示在这里
