Fetching the paper…
Reading the bibliography…
We propose Episodic Backward Update (EBU) - a novel deep reinforcement learning algorithm with a direct value propagation.
Watkins., C. J. C. H. Learning from delayed rewards. Ph.D. thesis, University of Cambridge England, 1989
1989
Earlier work this paper cites.
Lin, L-J. Programming Robots Using Reinforcement Learning and Teaching. In Association for the Advancement of Artificial Intelligence (AAAI), 781-786, 1991
1991
Earlier work this paper cites.
Lin, L-J. Self-improving reactive agents based on reinforcement learning, planning and teaching. Machine Learning, 293-321, 1992
1992
Earlier work this paper cites.
Watkins., C. J. C. H., and Dayan, P. Q-learning. Machine Learning, 272-292, 1992
1992
Earlier work this paper cites.
Bertsekas, D. P., and Tsitsiklis, J. N. Neuro-Dynamic Programming. Athena Scientific, 1996
1996
Earlier work this paper cites.
LeCun, Y., Bottou, L., Bengio, Y., and Haffner, P. Gradient-based learning applied to document recognition. In the Institute of Electrical and Electronics Engineers (IEEE), 86, 2278-2324, 1998
1998
Earlier work this paper cites.
Sutton, R. S., and Barto, A. G. Reinforcement Learning: An Introduction. MIT Press, 1998
1998
Earlier work this paper cites.
Melo, F. S. Convergence of Q-learning: A simple proof, Institute Of Systems and Robotics, Tech. Rep, 2001
2001
Earlier work this paper cites.
Lengyel, M., and Dayan, P. Hippocampal Contributions to Control: The Third Way. In Advances in Neural Information Processing Systems (NIPS), 889-896, 2007
2007
Cited alongside, same era.
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M. The arcade learning environment: An evaluation platform for general agents. Journal of Artificial Intelligence Research, 47:253-279, 2013
2013
Cited alongside, same era.
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., Petersen, S., Beattie, C., Sadik, A., Antonoglou, I., King, H., Kumaran, D., Wierstra, D., Legg, S., and Hassabis, D. Human-level control through deep reinforcement learning. Nature, 518(7540):529-533, 2015
2015
Cited alongside, same era.
Bellemare, M. G., Srinivasan, S., Ostrovski, G., Schaul, T., Saxton, D., and Munos, R. Unifying count-based exploration and intrinsic motivation. In Advances in Neural Information Processing Systems (NIPS), 1471-1479, 2016
Silver, D., Huang, A., Maddison C. J., Guez, A., Sifre, L., van den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., Dieleman, S., Grewe, D., Nham, J., Kalchbrenner, N., Sutskever, I., Lillicrap, T., Leach, M., Kavukcuoglu, K., Graepel, T., and Hassabis, D. Mastering the game of Go with deep neural networks and tree search. Nature, 529:484-489, 2016
2016
Later among the works it cites.
van Hasselt, H., Guez, A., and Silver, D. Deep Reinforcement Learning with Double Q-learning. In Association for the Advancement of Artificial Intelligence (AAAI), 2094-2100, 2016
2016
Later among the works it cites.
Wang, Z., Schaul, T., Hessel, M., van Hasselt, H., Lanctot, M., and de Freitas, N. Dueling Network Architectures for Deep Reinforcement Learning. In International Conference on Machine Learning (ICML), 1995-2003, 2016
2016
Later among the works it cites.
He, F. S., Liu, Y., Schwing, A. G., and Peng, J. Learning to play in a day: Faster deep reinforcement learning by optimality tightening. In International Conference on Learning Representations (ICLR), 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2016
Cited alongside, same era.
2016
Cited alongside, same era.
Harutyunyan, A., Bellemare, M. G., Stepleton, T., and Munos, R. Q( λ \lambda ) with off-policy corrections. In International Conference on Algorithmic Learning Theory (ALT), 305-320, 2016
2016
Cited alongside, same era.
Munos, R., Stepleton, T., Harutyunyan, A., and Bellemare, M. G. Safe and efficient off-policy reinforcement learning. In Advances in Neural Information Processing Systems (NIPS), 1046-1054, 2016
2016
Cited alongside, same era.
Schaul, T., Quan, J., Antonoglou, I., and Silver, D. Prioritized Experience Replay. In International Conference on Learning Representations (ICLR), 2016
2016
Cited alongside, same era.
2017
Later among the works it cites.
Pritzel, A., Uria, B., Srinivasan, S., Puig-’domenech, A., Vinyals, O., Hassabis, D., Wierstra, D., and Blundell, C. Neural Episodic Control. In International Conference on Machine Learning (ICML), 2827-2836, 2017
2017
Later among the works it cites.
2018
Closest in time.
Hansen, S., Pritzel, A., Sprechmann, P., Barreto, A., and Blundell, C. Fast deep reinforcement learning using online adjustments from the past. In Advances in Neural Information Processing Systems (NIPS), 10590–10600, 2018
2018
Closest in time.