ARTICLE
31 August 2026

Multi-Agent Reinforcement Learning for Dynamic Portfolio Allocation

Wanqing Feng1
Show Less
1 Rice Management Company, Houston, TX, USA
PBES 2026 , 9(8), 201–206; https://doi.org/10.18063/PBES.v9i8.15160
© 2026 by the Author(s). Licensee Whioce Publishing, Singapore. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution 4.0 International License ( https://creativecommons.org/licenses/by/4.0/ )
Abstract

The old portfolio model cannot be used to cope with the non-stationarity of the market, changes in asset correlation, and the pursuit of the risk-return trade-off. The problems above will be solved by applying dynamic portfolio allocation based on multi-agent reinforcement learning in this paper. The three special agents in the hierarchical heterogeneous multi-agent architecture are market perception, return prediction, and risk control; a centralized training with decentralized execution mode is employed, a graph attention network is used to obtain real-time inter-asset relationships, and thus dynamic iterative optimization of allocation weights and rebalancing strategies is achieved. Based on backtesting with CSI 300 constituents, 10-year treasury bond futures and the Nanhua Commodity Index, it was found that the proposed method can reduce the maximum drawdown and improve both the Sharpe ratio and excess return considerably compared to the traditional mean-variance model and the single-agent model; it has good market adaptability and risk resistance, and is thus a good and practical technical solution for the dynamic optimization of quantitative portfolios.

Keywords
Multi-agent reinforcement learning
Dynamic portfolio
Asset allocation
Risk control
Quantitative investment
References

[1] Zhang B, 2025, Graph Attention-Based Heterogeneous Multi-Agent Deep Reinforcement Learning for Adaptive Portfolio Optimization. Scientific Reports, 16(1): 2674.

[2] Shen Y, Li C, Scaillet O, et al., 2026, Dynamic Portfolio Allocation Under Market Incompleteness and Wealth Effects. Operations Research, 74(1): 93.

[3] Cousin A, Lelong J, Picard T, 2023, Mean-variance Dynamic Portfolio Allocation with Transaction Costs: A Wiener Chaos Expansion Approach. Applied Mathematical Finance, 30(6): 313–353.

[4] Siyang P, Shaojun G, Yonghong L, 2022, Large Dynamic Covariance Matrix Estimation with an Application to Portfolio Allocation: A Semiparametric Reproducing Kernel Hilbert Space Approach. Journal of Systems Science and Complexity, 35(4): 1429–1457.

[5] Siyang P, Shaojun G, Yonghong L, 2022, Large Dimensional Portfolio Allocation Based on a Mixed Frequency Dynamic Factor Model. Econometric Reviews, 41(5): 539–563.

[6] Haitao L, Chongfeng W, Chunyang Z, 2021, Time-Varying Risk Aversion and Dynamic Portfolio Allocation. Operations Research, 70(1).

[7] Jin X, Lehnert T, 2018, Large Portfolio Risk Management and Optimal Portfolio Allocation with Dynamic Elliptical Copulas. Dependence Modeling, 6(1): 19–46.

[8] Garcia CR, González V, Contreras J, et al., 2017, Applying Modern Portfolio Theory for a Dynamic Energy Portfolio Allocation in Electricity Markets. Electric Power Systems Research, 150: 11–23.

[9] Pinar CM, 2014, About the Dynamic Allocation of Robust Portfolio Against the Uncertainty of the Average Yields. INFOR: Information Systems and Operational Research, 52(1): 14–19.

[10] Thomaidis SN, Roumpis E, Kondakis N, 2010, Optimal Portfolio Allocation Strategies with Dynamic Factor Models. Int. J. of Financial Markets and Derivatives, 1(4): 352–370.

 

Share
Back to top