Fetching the paper…
Reading the bibliography…
The efficiency of reinforcement learning algorithms depends critically on a few meta-parameters that modulates the learning updates and the trade-off between exploration and exploitation.
Reducing bias and inefficiency in the selection algorithm
Baker, J. E. (1987) · 1987
Earlier work this paper cites.
Learning to predict by the method of temporal differences
Sutton, R. S. (1988) · 1988
Earlier work this paper cites.
Adapting bias by gradient descent: an incremental version of the delta-bar-delta
Sutton, R. S. (1992) · 1992
Earlier work this paper cites.
On-line Q-learning using connectionist systems
Rummery, G. A. and Niranjan, M. (1994) · 1994
Earlier work this paper cites.
Differentiation of learning abilities – a case study on optimizing parameter values in q-learning by a genetic algorithm
Unemi, T., Nagaoyoshi, M., Hirayama, N., Nade, T., Yano, K., and Masujima, Y. (1994) · 1994
Earlier work this paper cites.
Temporal differences based policy iteration and applications in neuro-dynamic programming
Bertsekas, D. P. and Ioffe, S. (1996) · 1996
Earlier work this paper cites.
Generalization in reinforcement learning: Successful examples using sparse coarse coding
Sutton, R. S. (1996) · 1996
Earlier work this paper cites.
How to lose at Tetris
Burgiel, H. (1997) · 1997
Earlier work this paper cites.
Reinforcement Learning: An Introduction
Sutton, R. S. and Barto, A. (1998) · 1998
Earlier work this paper cites.
Control of exploitation–exploration meta-parameter in reinforcement learning
Ishii, S., Yoshida, W., and Yoshimoto, J. (2002) · 2002
Cited alongside, same era.
Evolution of meta-parameters in reinforcement learning algorithm
Eriksson, A., Capi, G., and Doya, K. (2003) · 2003
Cited alongside, same era.
Meta-learning in reinforcement learning
Schweighofer, N. and Doya, K. (2003) · 2003
Cited alongside, same era.
Co-evolution of shaping rewards and meta-parameters in reinforcement learning
Elfwing, S., Uchibe, E., Doya, K., and Christensen, H. I. (2008) · 2008
Cited alongside, same era.
A meta-learning method based on temporal difference error
Kobayashi, K., Mizoue, H., Kuremoto, T., and Obayashi, M. (2009) · 2009
Cited alongside, same era.
Temporal difference bayesian model averaging: A bayesian perspective on adapting lambda
Downey, C. and Sanner, S. (2010) · 2010
Neural network ensembles in reinforcement learning
Faußer, S. and Schwenker, F. (2013) · 2013
Later among the works it cites.
Approximate dynamic programming finally performs well in the game of Tetris
Gabillon, V., Ghavamzadeh, M., and Scherrer, B. (2013) · 2013
Later among the works it cites.
How to discount deep reinforcement learning: Towards new dynamic strategies
François-Lavet, V., Fonteneau, R., and Ernst, D. (2015) · 2015
Later among the works it cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., Petersen, S., Beattie, C., Sadik, A., Antonoglou, I., King, H., Kumaran, D., Wierstra, D., Legg, S., and Hassabis, D. (2015) · 2015
Later among the works it cites.
Massively parallel methods for deep reinforcement learning
Nair, A., Srinivasan, P., Blackwell, S., Alcicek, C., Fearon, R., Maria, A. D., Panneershelvam, V., Suleyman, M., Beattie, C., Petersen, S., Legg, S., Mnih, V., Kavukcuoglu, K., and Silver, D. (2015) · 2015
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
SZ-Tetris as a benchmark for studying key problems of reinforcement learning
Szita, I. and Szepesvári, C. (2010) · 2010
Cited alongside, same era.
Darwinian embodied evolution of the learning ability for survival
Elfwing, S., Uchibe, E., Doya, K., and Christensen, H. I. (2011) · 2011
Cited alongside, same era.
The arcade learning environment: An evaluation platform for general agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M. (2013) · 2013
Cited alongside, same era.
Philosophie Zoologique
Lamarck, J. B. (1809)
Cited in the paper.
Later among the works it cites.
Deep reinforcement learning with double q-learning
van Hasselt, H., Guez, A., and Silver, D. (2015) · 2015
Later among the works it cites.
Adaptive λ \lambda least-squares temporal difference learning
Mann, T. A., Penedones, H., and Hester, T. (2016) · 2016
Later among the works it cites.
Sigmoid-weighted linear units for neural network function approximation in reinforcement learning
Elfwing, S., Uchibe, E., and Doya, K. (2017) · 2017
Closest in time.