Fetching the paper…
Reading the bibliography…
The field of reinforcement learning (RL) is facing increasingly challenging domains with combinatorial complexity.
Making the world differentiable: On using self-supervised fully recurrent neural networks for dynamic reinforcement learning and planning in non-stationary environments
Schmidhuber, J · 1990
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J · 1992
Earlier work this paper cites.
Long short-term memory
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
Neural network design for j function approximation in dynamic programming
Pang, X. and Werbos, P · 1998
Earlier work this paper cites.
Reinforcement learning: An introduction
Sutton, R. S., Barto, A. G., et al · 1998
Earlier work this paper cites.
Using abstraction for planning in sokoban
Botea, A., Müller, M., and Schaeffer, J · 2003
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Model regularization for stable sample rollouts
Talvitie, E · 2014
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Earlier work this paper cites.
Convolutional lstm network: A machine learning approach for precipitation nowcasting
Xingjian, S., Chen, Z., Wang, H., Yeung, D.-Y., Wong, W.-K., and Woo, W.-c · 2015
Earlier work this paper cites.
Adaptive computation time for recurrent neural networks
Graves, A · 2016
Earlier work this paper cites.
Highway and residual networks learn unrolled iterative estimation
Greff, K., Srivastava, R. K., and Schmidhuber, J · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Cited alongside, same era.
Reinforcement learning with unsupervised auxiliary tasks
Jaderberg, M., Mnih, V., Czarnecki, W. M., Schaul, T., Leibo, J. Z., Silver, D., and Kavukcuoglu, K · 2016
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., and Kavukcuoglu, K · 2016
Cited alongside, same era.
The predictron: End-to-end learning and planning
Silver, D., van Hasselt, H., Hessel, M., Schaul, T., Guez, A., Harley, T., Dulac-Arnold, G., Reichert, D., Rabinowitz, N., Barreto, A., et al · 2016
Cited alongside, same era.
Value iteration networks
Tamar, A., Wu, Y., Thomas, G., Levine, S., and Abbeel, P · 2016
Cited alongside, same era.
Imagination-augmented agents for deep reinforcement learning
Racanière, S., Weber, T., Reichert, D., Buesing, L., Guez, A., Rezende, D. J., Badia, A. P., Vinyals, O., Heess, N., Li, Y., et al · 2017
Later among the works it cites.
Lipschitz continuity in model-based reinforcement learning
Asadi, K., Misra, D., and Littman, M. L · 2018
Later among the works it cites.
Learning and querying fast generative models for reinforcement learning
Buesing, L., Weber, T., Racaniere, S., Eslami, S., Rezende, D., Reichert, D. P., Viola, F., Besse, F., Gregor, K., Hassabis, D., et al · 2018
Later among the works it cites.
Quantifying generalization in reinforcement learning
Cobbe, K., Klimov, O., Hesse, C., Kim, T., and Schulman, J · 2018
Later among the works it cites.
Visual foresight: Model-based deep reinforcement learning for vision-based robotic control
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zhang, C., Bengio, S., Hardt, M., Recht, B., and Vinyals, O · 2016
Cited alongside, same era.
A closer look at memorization in deep networks
Arpit, D., Jastrzębski, S., Ballas, N., Krueger, D., Bengio, E., Kanwal, M. S., Maharaj, T., Fischer, A., Courville, A., Bengio, Y., et al · 2017
Cited alongside, same era.
Treeqn and atreec: Differentiable tree planning for deep reinforcement learning
Farquhar, G., Rocktäschel, T., Igl, M., and Whiteson, S · 2017
Cited alongside, same era.
Darla: Improving zero-shot transfer in reinforcement learning
Higgins, I., Pal, A., Rusu, A., Matthey, L., Burgess, C., Pritzel, A., Botvinick, M., Blundell, C., and Lerchner, A · 2017
Cited alongside, same era.
Squeeze-and-excitation networks
Hu, J., Shen, L., and Sun, G · 2017
Cited alongside, same era.
Residual connections encourage iterative inference
Jastrzebski, S., Arpit, D., Ballas, N., Verma, V., Che, T., and Bengio, Y · 2017
Cited alongside, same era.
Value prediction network
Oh, J., Singh, S., and Lee, H · 2017
Cited alongside, same era.
Ebert, F., Finn, C., Dasari, S., Xie, A., Lee, A., and Levine, S · 2018
Later among the works it cites.
Impala: Scalable distributed deep-rl with importance weighted actor-learner architectures
Espeholt, L., Soyer, H., Munos, R., Simonyan, K., Mnih, V., Ward, T., Doron, Y., Firoiu, V., Harley, T., Dunning, I., et al · 2018
Later among the works it cites.
Learning to search with mctsnets
Guez, A., Weber, T., Antonoglou, I., Simonyan, K., Vinyals, O., Wierstra, D., Munos, R., and Silver, D · 2018
Later among the works it cites.
Distributed prioritized experience replay
Horgan, D., Quan, J., Budden, D., Barth-Maron, G., Hessel, M., Van Hasselt, H., and Silver, D · 2018
Later among the works it cites.
Lee, L., Parisotto, E., Chaplot, D. S., Xing, E., and Salakhutdinov, R · 2018
Later among the works it cites.
Single-agent policy tree search with guarantees
Orseau, L., Lelis, L., Lattimore, T., and Weber, T · 2018
Later among the works it cites.
Prefrontal cortex as a meta-reinforcement learning system
Wang, J. X., Kurth-Nelson, Z., Kumaran, D., Tirumala, D., Soyer, H., Leibo, J. Z., Hassabis, D., and Botvinick, M · 2018
Later among the works it cites.
Relational deep reinforcement learning
Zambaldi, V., Raposo, D., Santoro, A., Bapst, V., Li, Y., Babuschkin, I., Tuyls, K., Reichert, D., Lillicrap, T., Lockhart, E., et al · 2018
Later among the works it cites.