Volume 10,Issue 3
Power system operation optimization faces dual challenges from energy structure transformation and extreme environmental conditions. Traditional unit control methods demonstrate limitations in addressing renewable energy volatility, load demand uncertainty, and sudden system disturbances. Deep reinforcement learning, through constructing a state-action-reward decision framework, effectively handles the time-varying, nonlinear, and uncertain characteristics of complex systems, providing new technical pathways for unit operation optimization. Studies show that applications of voltage regulation frameworks based on gated Markov decision processes and reinforcement learning in optimizing high-pressure feedwater heater operations, along with the integration of Hooke-Jeeves algorithm and deep deterministic strategy gradient methods in air handling unit control, all validate deep reinforcement learning’s unique advantages in solving multi-objective optimization problems for power generation units.