An Off-Policy Reinforcement Learning-Based Adaptive Optimization Method for Dynamic Resource Allocation Problem.

He, Baiyang; Meng, Ying; Tang, Lixin · IEEE Trans Neural Netw Learn Syst · 2025

Where this comes from

Abstract

In this article, an adaptive optimization method is proposed for the dynamic resource allocation problem (RAP) with multiple objectives in the manufacturing industry. In the proposed method, a novel reinforcement learning method (DSAC-ERCE) is designed to adaptively set the weights for multiple objectives, and then the optimization method is adopted to generate the noninferior solutions in each time period. To ensure DSAC-ERCE's performance in dynamic and complex resource allocation environments, we develop a state-encoding network with a proposed information entropy attention mechanism to encode the state. Then, we introduce a new reward function to escape from the local optima of the policy and further present a conditional entropy policy to enhance the policy network. In addition, we demonstrate the feasibility of improving the quality of actions and present a boundary method for high-quality actions. We also introduce an optimization model to automatically adjust the temperature parameter in DSAC-ERCE. Furthermore, we compare and analyze our approach with other state-of-the-art reinforcement learning methods. The experiments illustrate that DSAC-ERCE outperforms state-of-the-art reinforcement learning methods. Moreover, DSAC-ERCE can be generalized to solve optimization problems with two to five objectives, problems with linear, quadratic, cubic, logarithmic, or inverse objectives, and problems with diverse structures.