Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Diffusion policies have demonstrated impressive performance in robot control, yet are difficult to improve online when 0-shot performance isn’t enough. To address this challenge, we introduce DSRL: Diffusion Steering via Reinforcement Learning. (1/n)

65,162 görüntüleme • 1 yıl önce •via X (Twitter)

25 Yorum

Andrew Wagenmaker profil fotoğrafı
Andrew Wagenmaker1 yıl önce

DSRL trains a lightweight policy via RL to select input noise to the diffusion policy's denoising process, steering it to desired behaviors. We find this leads to very sample-efficient improvement, and avoids challenges typically encountered with diffusion policies + RL. (2/n)

Andrew Wagenmaker profil fotoğrafı
Andrew Wagenmaker1 yıl önce

DSRL enables efficient improvement of real-world diffusion policies for robotic control. We apply it to several different robot embodiments, and find that it is able to improve performance from <30% success to >90% in anywhere from 30-60 minutes of online training. (3/n)

Andrew Wagenmaker profil fotoğrafı
Andrew Wagenmaker1 yıl önce

We also apply DSRL to state-of-the-art flow-based generalist policies, in particular pi0 from Physical Intelligence. DSRL is able to improve pi0 in real-world deployment, on some tasks taking success from 25% to 90% in <90 minutes of online training. (4/n)

Andrew Wagenmaker profil fotoğrafı
Andrew Wagenmaker1 yıl önce

Uncut training timelapse of DSRL on WidowX pick-and-place task. (5/n)

Andrew Wagenmaker profil fotoğrafı
Andrew Wagenmaker1 yıl önce

Real-world DSRL training is highly stable. As any "action" played by the noise-space RL policy is just initial noise for the denoising process, even early in training the denoised actions look like actions from a BC-trained policy, rather than an unconverged RL policy. (6/n)

Andrew Wagenmaker profil fotoğrafı
Andrew Wagenmaker1 yıl önce

DSRL can, in principle, be instantiated with any RL algorithm. We propose, however, a SAC-based variant, Noise-Aliased DSRL, that takes advantage of a diffusion policy's tendency to map different noise to similar actions, and allows for fully off-policy training. (7/n)

Andrew Wagenmaker profil fotoğrafı
Andrew Wagenmaker1 yıl önce

Noise-Aliased DSRL trains two Q-functions, one on the original action space via TD learning, and one on the noise action space by distilling the first Q function. This allows learning from offline data and improves online sample efficiency by as much as 2x. (8/n)

Andrew Wagenmaker profil fotoğrafı
Andrew Wagenmaker1 yıl önce

In simulation, we find that DSRL substantially outperforms all existing approaches to improving diffusion policies online on benchmarks such as Robomimic. (9/n)

Andrew Wagenmaker profil fotoğrafı
Andrew Wagenmaker1 yıl önce

DSRL is also a competitive offline RL procedure: we find that first training a diffusion/flow policy on an offline dataset, then applying DSRL to steer it to high-reward behavior using the offline data performs on par with state-of-the-art offline RL methods on OGBench. (10/n)

Andrew Wagenmaker profil fotoğrafı
Andrew Wagenmaker1 yıl önce

Fun collaboration with @mitsuhiko_nm, @yunchuzh, @seohong_park, Waleed Yagoub, Anusha Nagabandi, @abhishekunique7, @svlevine! (11/n) Please see the paper and website for additional results! Website: Paper:

Stone Tao profil fotoğrafı
Stone Tao1 yıl önce

looks awesome and the presentation video is extremely clear! I was looking at the real world training and felt that there isn’t much randomization of the start state eg pick/place the cube into the bowl, cube is always fixed. Have you tried larger reset regions? What’s the sample efficiency in those scenarios?

Andrew Wagenmaker profil fotoğrafı
Andrew Wagenmaker1 yıl önce

We've played around with this a bit, and it still works, it just takes somewhat longer to train (the exact amount probably depends how diverse initial state is, but didn't seem to be significantly more from what we saw).

Siddarth profil fotoğrafı
Siddarth1 yıl önce

Congratulations on your work Andrew! This is actually highly related to our work (just presented at ICML) Outsourced Diffusion Sampling. In our work we learn to sample from high value tilted distributions from generative model priors by learning a policy over the latent space. We actually parameterize that policy itself as a diffusion model, trained with off-policy RL. It is suitable for offline RL too, and naturally enforces conservatism because it samples from the Bayesian posterior p_{bc}(s | a)exp(beta*Q(s, a))

Vaishak Kumar profil fotoğrafı
Vaishak Kumar1 yıl önce

Looks like the link to the code is broken

Andrew Wagenmaker profil fotoğrafı
Andrew Wagenmaker1 yıl önce

We will post code soon!

Francesco Capuano profil fotoğrafı
Francesco Capuano1 yıl önce

@AdilZtn ✨👀

Aayush Karan profil fotoğrafı
Aayush Karan1 yıl önce

Really cool work! In case you're interested, we found that latent noises for diffusion corresponding to high sample reward tend to live in these disjoint "balls" in noise space, and that this same noise initialization trick works really well for steering images too!

Andrew Wagenmaker profil fotoğrafı
Andrew Wagenmaker1 yıl önce

Thanks for the pointer! Will definitely take a look at this

Muyu He profil fotoğrafı
Muyu He1 yıl önce

Interesting to see that rewards converge early on most RL tasks. Does the selection of the initial noise follow the same pattern in the latent space or is there a delayed effect?

Brian Smith profil fotoğrafı
Brian Smith1 yıl önce

Very cool work! Are you planning on publishing your code? I had a question about how you did real-world policy rollouts, but I saw the code button on the website links to the website itself.

Andrew Wagenmaker profil fotoğrafı
Andrew Wagenmaker1 yıl önce

Yes, we do plan to publish code, hopefully in the near future!

Dan profil fotoğrafı
Dan1 yıl önce

Your post is fascinating, especially for someone like me who’s always looking to optimise. It’s just clicked for me that diffusion is right at the heart of why robots move in that ‘robotic’ way and it’s directly tied to the whole step-by-step refinement process.

Ted Xiao profil fotoğrafı
Ted Xiao1 yıl önce

Congratulations, cool work! Do you find that the space of diffusion noise is well-organized? Do different types of noise map to semantic or interpretable behavior modes?

Andrew Wagenmaker profil fotoğrafı
Andrew Wagenmaker1 yıl önce

We ran some very preliminary experiments exploring this, and our results suggested that the noise space does exhibit some semantic structure. This really requires further investigation though, and would be a cool direction for future work!

Ted Xiao profil fotoğrafı
Ted Xiao1 yıl önce

Thanks for the insights, congrats again!

Benzer Videolar

MaterialFusion Enhancing Inverse Rendering with Material Diffusion Priors discuss: Recent works in inverse rendering have shown promise in using multi-view images of an object to recover shape, albedo, and materials. However, the recovered components often fail to render accurately under new lighting conditions due to the intrinsic challenge of disentangling albedo and material properties from input images. To address this challenge, we introduce MaterialFusion, an enhanced conventional 3D inverse rendering pipeline that incorporates a 2D prior on texture and material properties. We present StableMaterial, a 2D diffusion model prior that refines multi-lit data to estimate the most likely albedo and material from given input appearances. This model is trained on albedo, material, and relit image data derived from a curated dataset of approximately ~12K artist-designed synthetic Blender objects called BlenderVault. we incorporate this diffusion prior with an inverse rendering framework where we use score distillation sampling (SDS) to guide the optimization of the albedo and materials, improving relighting performance in comparison with previous work. We validate MaterialFusion's relighting performance on 4 datasets of synthetic and real objects under diverse illumination conditions, showing our diffusion-aided approach significantly improves the appearance of reconstructed objects under novel lighting conditions. We intend to publicly release our BlenderVault dataset to support further research in this field.

AK

22,959 görüntüleme • 2 yıl önce