ๆญฃๅจๅ ่ฝฝ่ง้ข...
่ง้ขๅ ่ฝฝๅคฑ่ดฅ
๐ค How to fine-tune an Imitation Learning policy (e.g., Diffusion Policy, ACT) with RL? As an RL practitioner, Iโve been struggling with this problem for a while. Hereโs why itโs tough: 1๏ธโฃ Special designs (usually for multimodal action distributions) in modern IL models make them non-trivial to fine-tune by... show more
17,018 ๆฌก่ง็ โข 1 ๅนดๅ โขvia X (Twitter)
0 ๆก่ฏ่ฎบ
ๆๆ ่ฏ่ฎบ
ๅๅงๅธๅญ็่ฏ่ฎบๅฐๆพ็คบๅจ่ฟ้
