Diffusion policy distillation for offline reinforcement learning.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 40532527.
- Also identified by DOI 10.1016/j.neunet.2025.107694.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Offline reinforcement learning aims to learn a well-performing target policy from a static empirical dataset. Leveraging its powerful distribution expression capabilities, the diffusion model has been widely adopted as a type of policy in offline reinforcement learning. However, sampling a single action from diffusion policy necessitates a multi-step denoise process, which results in slow decision-making speed and poses challenges for application in real-time control tasks. Inspired by the teacher-student mechanism in human learning, this paper proposes a diffusion policy distillation (DPD) framework, which employs a deterministic policy to distill the target policy induced by the diffusion model. Although the deterministic policy cannot express the complex behavior policy induced by the empirical dataset properly, it can effectively learn a relevant target policy. Moreover, since the distillated deterministic policy is one-step, it avoids the need for iterative denoising, thereby inheriting the performance of the target policy while effectively improving the decision-making speed. DPD is plug-and-play and thus can be combined with offline reinforcement learning methods based on diffusion policy. Experimental results on D4RL Gym-MuJoCo datasets indicate that the distillation policy can achieve a higher normalized score than the original policy with a lower standard deviation, and improve the decision-making speed by over 10 times.
Medical subject headings
- Reinforcement, Psychology
- Neural Networks, Computer