Fetching the paper…
Reading the bibliography…
Modern reinforcement learning algorithms reach super-human performance on many board and video games, but they are sample inefficient, i.e.
Introduction and removal of reward, and maze performance in rats
Tolman, E. C. and Honzik, C. H · 1930
Earlier work this paper cites.
Global optimization of a neural network-hidden markov model hybrid
Bengio, Y., De Mori, R., Flammia, G., and Kompe, R · 1992
Earlier work this paper cites.
Improving generalization for temporal difference learning: The successor representation
Dayan, P · 1993
Earlier work this paper cites.
Prioritized sweeping: Reinforcement learning with less data and less time
Moore, A. W. and Atkeson, C. G · 1993
Earlier work this paper cites.
Efficient learning and planning within the dyna framework
Peng, J. and Williams, R. J · 1993
Earlier work this paper cites.
Similarity estimation techniques from rounding algorithms
Charikar, M. S · 2002
Earlier work this paper cites.
Neural fitted q iteration–first experiences with a data efficient neural reinforcement learning method
Riedmiller, M · 2005
Earlier work this paper cites.
Towards a unified theory of state abstraction for MDPs
Li, L., Walsh, T. J., and Littman, M. L · 2006
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E · 2012
Earlier work this paper cites.
The Arcade Learning Environment: An evaluation platform for general agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M · 2013
Earlier work this paper cites.
Auto-encoding variational bayes
Kingma, D. P. and Welling, M · 2013
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M · 2013
Earlier work this paper cites.
Efficient planning in MDPs by small backups
Van Seijen, H. and Sutton, R. S · 2013
Earlier work this paper cites.
Generative adversarial nets
Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y · 2014
Cited alongside, same era.
Neural variational inference and learning in belief networks
Mnih, A. and Gregor, K · 2014
Cited alongside, same era.
Stochastic backpropagation and approximate inference in deep generative models
Rezende, D. J., Mohamed, S., and Wierstra, D · 2014
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Cited alongside, same era.
Schaul, T., Quan, J., Antonoglou, I., and Silver, D · 2015
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., and Kavukcuoglu, K · 2016
Later among the works it cites.
Scaling factorial hidden markov models: Stochastic variational inference without messages
Ng, Y. C., Chilinski, P. M., and Silva, R · 2016
Later among the works it cites.
Is prioritized sweeping the better episodic control?
Brea, J · 2017
Later among the works it cites.
TreeQN and ATreeC: Differentiable Tree Planning for Deep Reinforcement Learning
Farquhar, G., Rocktäschel, T., Igl, M., and Whiteson, S · 2017
Later among the works it cites.
Neural Network Dynamics for Model-Based Deep Reinforcement Learning with Model-Free Fine-Tuning
Nagabandi, A., Kahn, G., Fearing, R. S., and Levine, S · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Blundell, C., Uria, B., Pritzel, A., Li, Y., Ruderman, A., Leibo, J. Z., Rae, J., Wierstra, D., and Hassabis, D · 2016
Cited alongside, same era.
Reinforcement learning with unsupervised auxiliary tasks
Jaderberg, M., Mnih, V., Czarnecki, W. M., Schaul, T., Leibo, J. Z., Silver, D., and Kavukcuoglu, K · 2016
Cited alongside, same era.
Categorical reparameterization with gumbel-softmax
Jang, E., Gu, S., and Poole, B · 2016
Cited alongside, same era.
ViZDoom: A Doom-based AI research platform for visual reinforcement learning
Kempka, M., Wydmuch, M., Runc, G., Toczek, J., and Jaśkowski, W · 2016
Cited alongside, same era.
Deep successor reinforcement learning
Kulkarni, T. D., Saeedi, A., Gautam, S., and Gershman, S. J · 2016
Cited alongside, same era.
The Concrete Distribution: A Continuous Relaxation of Discrete Random Variables
Maddison, C. J., Mnih, A., and Whye Teh, Y · 2016
Cited alongside, same era.
Learning to Navigate in Complex Environments
Mirowski, P., Pascanu, R., Viola, F., Soyer, H., Ballard, A. J., Banino, A., Denil, M., Goroshin, R., Sifre, L., Kavukcuoglu, K., Kumaran, D., and Hadsell, R · 2016
Cited alongside, same era.
Oh, J., Singh, S., and Lee, H · 2017
Later among the works it cites.
Neural episodic control
Pritzel, A., Uria, B., Srinivasan, S., Badia, A. P., Vinyals, O., Hassabis, D., Wierstra, D., and Blundell, C · 2017
Later among the works it cites.
Imagination-augmented agents for deep reinforcement learning
Racanière, S., Weber, T., Reichert, D., Buesing, L., Guez, A., Rezende, D. J., Badia, A. P., Vinyals, O., Heess, N., Li, Y., et al · 2017
Later among the works it cites.
Predictive representations can link model-based reinforcement learning to model-free mechanisms
Russek, E. M., Momennejad, I., Botvinick, M. M., Gershman, S. J., and Daw, N. D · 2017
Later among the works it cites.
Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm
Silver, D., Hubert, T., Schrittwieser, J., Antonoglou, I., Lai, M., Guez, A., Lanctot, M., Sifre, L., Kumaran, D., Graepel, T., Lillicrap, T., Simonyan, K., and Hassabis, D · 2017
Later among the works it cites.
The predictron: End-to-end learning and planning
Silver, D., van Hasselt, H., Hessel, M., Schaul, T., Guez, A., Harley, T., Dulac-Arnold, G., Reichert, D., Rabinowitz, N., Barreto, A., and Degris, T · 2017
Later among the works it cites.
REBAR: Low-variance, unbiased gradient estimates for discrete latent variable models
Tucker, G., Mnih, A., Maddison, C. J., Lawson, D., and Sohl-Dickstein, J · 2017
Later among the works it cites.
Reinforcement Learning: An Introduction
Sutton, R. S. and Barto, A. G · 2018
Closest in time.