Fetching the paper…
Reading the bibliography…
Natural policy gradient (NPG) methods with entropy regularization achieve impressive empirical success in reinforcement learning problems with large state-action spaces.
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu
1937
Earlier work this paper cites.
M. D. Donsker and S. S. Varadhan
1983
Earlier work this paper cites.
R. S. Sutton
1988
Earlier work this paper cites.
R. J. Williams
1992
Earlier work this paper cites.
D. P. Bertsekas and J. N. Tsitsiklis
1996
Earlier work this paper cites.
J. N. Tsitsiklis and B. Van Roy
1997
Earlier work this paper cites.
S.-I. Amari
1998
Earlier work this paper cites.
R. S. Sutton, D. A. McAllester, S. P. Singh, Y. Mansour, et al
1999
Earlier work this paper cites.
V. R. Konda and J. N. Tsitsiklis
2000
Earlier work this paper cites.
S. M. Kakade
2001
Earlier work this paper cites.
S. Bhatnagar, M. Ghavamzadeh, M. Lee, and R. S. Sutton
2007
Earlier work this paper cites.
A. Rahimi, B. Recht, et al
2007
Earlier work this paper cites.
J. Peters and S. Schaal
2008
Earlier work this paper cites.
S. Bhatnagar, R. S. Sutton, M. Ghavamzadeh, and M. Lee
2009
Earlier work this paper cites.
C. Szepesvári
2010
Earlier work this paper cites.
D. Hsu, S. M. Kakade, and T. Zhang
2012
Earlier work this paper cites.
B. Scherrer
2014
Cited alongside, same era.
S. Shalev-Shwartz and S. Ben-David
2014
Cited alongside, same era.
J. Schulman, S. Levine, P. Abbeel, M. Jordan, and P. Moritz
2015
Cited alongside, same era.
Y. Duan, X. Chen, R. Houthooft, J. Schulman, and P. Abbeel
2016
Cited alongside, same era.
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. Van Den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, et al
2016
Cited alongside, same era.
O. Nachum, M. Norouzi, K. Xu, and D. Schuurmans
2017
Cited alongside, same era.
M. Geist, B. Scherrer, and O. Pietquin
2019
Later among the works it cites.
B. Liu, Q. Cai, Z. Yang, and Z. Wang
2019
Later among the works it cites.
L. Wang, Q. Cai, Z. Yang, and Z. Wang
2019
Later among the works it cites.
S. Cen, C. Cheng, Y. Chen, Y. Wei, and Y. Chi
2020
Later among the works it cites.
J. Fan, Z. Wang, Y. Xie, and Z. Yang
2020
Later among the works it cites.
J. Martens
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2017
Cited alongside, same era.
G. Neu, A. Jonsson, and V. Gómez
2017
Cited alongside, same era.
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov
2017
Cited alongside, same era.
J. Bhandari, D. Russo, and R. Singal
2018
Cited alongside, same era.
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine
2018
Cited alongside, same era.
A. Jacot, F. Gabriel, and C. Hongler
2018
Cited alongside, same era.
J. Mei, C. Xiao, C. Szepesvari, and D. Schuurmans
2020
Later among the works it cites.
S. Satpathi, H. Gupta, S. Liang, and R. Srikant
2020
Later among the works it cites.
L. Shani, Y. Efroni, and S. Mannor
2020
Later among the works it cites.
2020
Later among the works it cites.
A. Agarwal, S. M. Kakade, J. D. Lee, and G. Mahajan
2021
Closest in time.
S. Khodadadian, P. R. Jhunjhunwala, S. M. Varma, and S. T. Maguluri
2021
Closest in time.
2022
Closest in time.
R. Yuan, S. S. Du, R. M. Gower, A. Lazaric, and L. Xiao
2022
Closest in time.
F. Hellström, G. Durisi, B. Guedj, and M. Raginsky
2023
Closest in time.