Fetching the paper…
Reading the bibliography…
In recent years, neural networks have enjoyed a renaissance as function approximators in reinforcement learning.
Asynchronous methods for deep reinforcement learning
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T. P., Harley, T., Silver, D., and Kavukcuoglu, K. (2016) · 1937
Earlier work this paper cites.
Information processing in dynamical systems: Foundations of harmony theory
Smolensky, P. (1986) · 1986
Earlier work this paper cites.
Learning to predict by the method of temporal differences
Sutton, R. S. (1988) · 1988
Earlier work this paper cites.
Unsupervised learning of distributions on binary vectors using two layer networks
Freund, Y. and Haussler, D. (1992) · 1992
Earlier work this paper cites.
Issues in using function approximation for reinforcement learning
Thrun, S. and Schwartz, A. (1993) · 1993
Earlier work this paper cites.
On-line Q-learning using connectionist systems
Rummery, G. A. and Niranjan, M. (1994) · 1994
Earlier work this paper cites.
Td-gammon, a self-teaching backgammon program, achieves master-level play
Tesauro, G. (1994) · 1994
Earlier work this paper cites.
Temporal differences based policy iteration and applications in neuro-dynamic programming
Bertsekas, D. P. and Ioffe, S. (1996) · 1996
Earlier work this paper cites.
Generalization in reinforcement learning: Successful examples using sparse coarse coding
Sutton, R. S. (1996) · 1996
Earlier work this paper cites.
How to lose at Tetris
Burgiel, H. (1997) · 1997
Earlier work this paper cites.
Reinforcement Learning: An Introduction
Sutton, R. S. and Barto, A. (1998) · 1998
Cited alongside, same era.
Digital selection and analogue amplification coexist in a cortex-inspired silicon circuit
Hahnloser, R. H. R., Sarpeshka, R., Mahowald, M. A., Douglas, R. J., and Seung, H. S. (2000) · 2000
Cited alongside, same era.
Training products of experts by minimizing contrastive divergence
Hinton, G. E. (2002) · 2002
Cited alongside, same era.
Dueling network architectures for deep reinforcement learning
Wang, Z., Schaul, T., Hessel, M., van Hasselt, H., Lanctot, M., and de Freitas, N. (2016) · 2003
Cited alongside, same era.
Improvements on learning tetris with cross entropy
Thiery, C. and Scherrer, B. (2009) · 2009
Cited alongside, same era.
SZ-Tetris as a benchmark for studying key problems of reinforcement learning
Szita, I. and Szepesvári, C. (2010) · 2010
Expected energy-based restricted boltzmann machine for classification
Elfwing, S., Uchibe, E., and Doya, K. (2015) · 2015
Later among the works it cites.
High-dimensional function approximation for knowledge-free reinforcement learning: a case study in SZ-Tetris
Jaskowski, W., Szubert, M. G., Liskowski, P., and Krawiec, K. (2015) · 2015
Later among the works it cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., Petersen, S., Beattie, C., Sadik, A., Antonoglou, I., King, H., Kumaran, D., Wierstra, D., Legg, S., and Hassabis, D. (2015) · 2015
Later among the works it cites.
Massively parallel methods for deep reinforcement learning
Nair, A., Srinivasan, P., Blackwell, S., Alcicek, C., Fearon, R., Maria, A. D., Panneershelvam, V., Suleyman, M., Beattie, C., Petersen, S., Legg, S., Mnih, V., Kavukcuoglu, K., and Silver, D. (2015) · 2015
Later among the works it cites.
Approximate modified policy iteration and its application to the game of tetris
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Double q-learning
van Hasselt, H. (2010) · 2010
Cited alongside, same era.
The arcade learning environment: An evaluation platform for general agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M. (2013) · 2013
Cited alongside, same era.
Neural network ensembles in reinforcement learning
Faußer, S. and Schwenker, F. (2013) · 2013
Cited alongside, same era.
Approximate dynamic programming finally performs well in the game of tetris
Gabillon, V., Ghavamzadeh, M., and Scherrer, B. (2013) · 2013
Cited alongside, same era.
Scherrer, B., Ghavamzadeh, M., Gabillon, V., Lesner, B., and Geist, M. (2015) · 2015
Later among the works it cites.
Deep reinforcement learning with double q-learning
van Hasselt, H., Guez, A., and Silver, D. (2015) · 2015
Later among the works it cites.
From free energy to expected energy: Improving energy-based value function approximation in reinforcement learning
Elfwing, S., Uchibe, E., and Doya, K. (2016) · 2016
Later among the works it cites.
Prioritized experience replay
Schaul, T., Quan, J., Antonoglou, I., and Silver, D. (2016) · 2016
Later among the works it cites.
Tetris AI, computer plays tetris
Fahey, C. (2003) · 2017
Closest in time.