Fetching the paper…
Reading the bibliography…
The combination of Monte-Carlo tree search (MCTS) with deep reinforcement learning has led to significant advances in artificial intelligence.
A theory of regularized markov decision processes
Geist, M., Scherrer, B., and Pietquin, O. (2019) · 1901
Earlier work this paper cites.
Discretizing continuous action space for on-policy optimization
Tang, Y. and Agrawal, S. (2019) · 1901
Earlier work this paper cites.
Policy gradient search: Online planning and expert iteration without search trees
Anthony, T., Nishihara, R., Moritz, P., Salimans, T., and Schulman, J. (2019) · 1904
Earlier work this paper cites.
V-MPO: On-policy maximum a posteriori policy optimization for discrete and continuous control
Song, H. F., Abdolmaleki, A., Springenberg, J. T., Clark, A., Soyer, H., Rae, J. W., Noury, S., Ahuja, A., Liu, S., Tirumala, D., et al. (2019) · 1909
Earlier work this paper cites.
Mastering Atari, go, chess and shogi by planning with a learned model
Schrittwieser, J., Antonoglou, I., Hubert, T., Simonyan, K., Sifre, L., Schmitt, S., Guez, A., Lockhart, E., Hassabis, D., Graepel, T., et al. (2019) · 1911
Earlier work this paper cites.
Eine informationstheoretische ungleichung und ihre anwendung auf beweis der ergodizitaet von markoffschen ketten
Csiszár, I. (1964) · 1964
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J. (1992) · 1992
Earlier work this paper cites.
Reinforcement Learning: An Introduction
Sutton, R. S. and Barto, A. G. (1998) · 1998
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S., McAllester, D. A., Singh, S. P., and Mansour, Y. (2000) · 2000
Earlier work this paper cites.
Q-learning in enormous action spaces via amortized approximate maximization
Van de Wiele, T., Warde-Farley, D., Mnih, A., and Mnih, V. (2020) · 2001
Earlier work this paper cites.
Using confidence bounds for exploitation-exploration trade-offs
Auer, P. (2002) · 2002
Earlier work this paper cites.
Convex optimization
Boyd, S. and Vandenberghe, L. (2004) · 2004
Earlier work this paper cites.
Bandit based Monte-Carlo planning
Kocsis, L. and Szepesvári, C. (2006) · 2006
Earlier work this paper cites.
On divergences and informations in statistics and information theory
Liese, F. and Vajda, I. (2006) · 2006
Earlier work this paper cites.
Modeling Purposeful Adaptive Behavior with the Principle of Maximum Causal Entropy
Ziebart, B. D. (2010) · 2010
Cited alongside, same era.
Multi-armed bandits with episode context
Rosin, C. D. (2011) · 2011
Cited alongside, same era.
A survey of monte carlo tree search methods
Browne, C. B., Powley, E., Whitehouse, D., Lucas, S. M., Cowling, P. I., Rohlfshagen, P., Tavener, S., Perez, D., Samothrakis, S., and Colton, S. (2012) · 2012
Cited alongside, same era.
Regret analysis of stochastic and nonstochastic multi-armed bandit problems
Bubeck, S., Cesa-Bianchi, N., et al. (2012) · 2012
Cited alongside, same era.
The arcade learning environment: An evaluation platform for general agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M. (2013) · 2013
Cited alongside, same era.
Discrete sequential prediction of continuous actions for deep rl
Metz, L., Ibarz, J., Jaitly, N., and Davidson, J. (2017) · 2017
Later among the works it cites.
A unified view of entropy-regularized Markov decision processes
Neu, G., Jonsson, A., and Gómez, V. (2017) · 2017
Later among the works it cites.
Value prediction network
Oh, J., Singh, S., and Lee, H. (2017) · 2017
Later among the works it cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O. (2017) · 2017
Later among the works it cites.
Maximum a posteriori policy optimisation
Abdolmaleki, A., Springenberg, J. T., Tassa, Y., Munos, R., Heess, N., and Riedmiller, M. (2018) · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dulac-Arnold, G., Evans, R., van Hasselt, H., Sunehag, P., Lillicrap, T., Hunt, J., Mann, T., Weber, T., Degris, T., and Coppin, B. (2015) · 2015
Cited alongside, same era.
Taming the noise in reinforcement learning via soft updates
Fox, R., Pakman, A., and Tishby, N. (2015) · 2015
Cited alongside, same era.
Trust region policy optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M., and Moritz, P. (2015) · 2015
Cited alongside, same era.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J. (2016) · 2016
Cited alongside, same era.
Combining policy gradient and Q-learning
O’Donoghue, B., Munos, R., Kavukcuoglu, K., and Mnih, V. (2016) · 2016
Cited alongside, same era.
Mastering the game of go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al. (2016) · 2016
Cited alongside, same era.
Openai baselines
Dhariwal, P., Hesse, C., Klimov, O., Nichol, A., Plappert, M., Radford, A., Schulman, J., Sidor, S., Wu, Y., and Zhokhov, P. (2017) · 2017
Cited alongside, same era.
Later among the works it cites.
Distributional policy gradients
Barth-Maron, G., Hoffman, M. W., Budden, D., Dabney, W., Horgan, D., TB, D., Muldal, A., Heess, N., and Lillicrap, T. (2018) · 2018
Later among the works it cites.
Learning to search with mctsnets
Guez, A., Weber, T., Antonoglou, I., Simonyan, K., Vinyals, O., Wierstra, D., Munos, R., and Silver, D. (2018) · 2018
Later among the works it cites.
Distributed prioritized experience replay
Horgan, D., Quan, J., Budden, D., Barth-Maron, G., Hessel, M., van Hasselt, H., and Silver, D. (2018) · 2018
Later among the works it cites.
Reinforcement learning and control as probabilistic inference: Tutorial and review
Levine, S. (2018) · 2018
Later among the works it cites.
Tassa, Y., Doron, Y., Muldal, A., Erez, T., Li, Y., Casas, D. d. L., Budden, D., Abdolmaleki, A., Merel, J., Lefrancq, A., et al. (2018) · 2018
Later among the works it cites.
Planning in entropy-regularized Markov decision processes and games
Grill, J.-B., Domingues, O. D., Ménard, P., Munos, R., and Valko, M. (2019) · 2019
Later among the works it cites.
Learning dexterous in-hand manipulation
Andrychowicz, O. M., Baker, B., Chociej, M., Jozefowicz, R., McGrew, B., Pachocki, J., Petron, A., Plappert, M., Powell, G., Ray, A., et al. (2020) · 2020
Closest in time.
Cloud TPU — Google Cloud
Google (2020) · 2020
Closest in time.
Combining Q-learning and search with amortized value estimates
Hamrick, J. B., Bapst, V., Sanchez-Gonzalez, A., Pfaff, T., Weber, T., Buesing, L., and Battaglia, P. W. (2020) · 2020
Closest in time.