通知公告

上海金融智能工程技术研究中心张立文教授团队发表CCF-A类人工智能顶会论文一篇

发布于:2026-07-14 10:40:26     浏览量:{动态访问次数}

发表日期:2023年7月23日 

论文名称:Variance control for distributional reinforcement learning

作者:Q. Kuang, Z. Zhu, L. Zhang, & F. Zhou

摘要:Although distributional reinforcement learning (DRL) has been widely examined in the past few years, very few studies investigate the validity of the obtained Q-function estimator in the distributional setting. To fully understand how the approximation errors of the Q-function affect the whole training process, we do some error analysis and theoretically show how to reduce both the bias and the variance of the error terms. With this new understanding, we construct a new estimator Quantiled Expansion Mean (QEM) and introduce a new DRL algorithm (QEMRL) from the statistical perspective. We extensively evaluate our QEMRL algorithm on a variety of Atari and Mujoco benchmark tasks and demonstrate that QEMRL achieves significant improvement over baseline algorithms in terms of sample efficiency and convergence performance.