Diffusion policy distillation for offline reinforcement learning.

Zhang, Jiazhi; Cheng, Yuhu; Chen, C L Philip; Zhang, Hengrui; Wang, Xuesong · Neural Netw · 2025

basic_science · Level V

Where this comes from

Abstract

Offline reinforcement learning aims to learn a well-performing target policy from a static empirical dataset. Leveraging its powerful distribution expression capabilities, the diffusion model has been widely adopted as a type of policy in offline reinforcement learning. However, sampling a single action from diffusion policy necessitates a multi-step denoise process, which results in slow decision-making speed and poses challenges for application in real-time control tasks. Inspired by the teacher-student mechanism in human learning, this paper proposes a diffusion policy distillation (DPD) framework, which employs a deterministic policy to distill the target policy induced by the diffusion model. Although the deterministic policy cannot express the complex behavior policy induced by the empirical dataset properly, it can effectively learn a relevant target policy. Moreover, since the distillated deterministic policy is one-step, it avoids the need for iterative denoising, thereby inheriting the performance of the target policy while effectively improving the decision-making speed. DPD is plug-and-play and thus can be combined with offline reinforcement learning methods based on diffusion policy. Experimental results on D4RL Gym-MuJoCo datasets indicate that the distillation policy can achieve a higher normalized score than the original policy with a lower standard deviation, and improve the decision-making speed by over 10 times.

Medical subject headings