Fetching the paper…
Reading the bibliography…
The breakthrough of deep Q-Learning on different types of environments revolutionized the algorithmic design of Reinforcement Learning to introduce more stable and robust algorithms, to that end many extensions to deep Q-Learning algorithm have been proposed to reduce the variance of the target values and the overestimation phenomena.
Bellman, Richard. A Markovian decision process. Indiana Univ. Math. J., 6:679–684, 1957
1957
Earlier work this paper cites.
Learning and sequential decision making AG Barto, R. S. Sutton, CJCH Watkins - Learning and computational neuroscience, 1989
1989
Earlier work this paper cites.
C. J. C. H.Watkins, Learning from Delayed Rewards. PhD thesis, King’s College, Cambridge, England, 1989
1989
Earlier work this paper cites.
Watkins, Christopher JCH and Dayan, Peter. Q-learning. Machine Learning, 8(3-4):279–292, 1992
1992
Earlier work this paper cites.
Thrun, Sebastian and Schwartz, Anton. Issues in using function approximation for reinforcement learning. In Proceedings of the 1993 Connectionist Models Summer School Hillsdale, NJ. Lawrence Erlbaum, 1993
1993
Earlier work this paper cites.
Rummery, Gavin A and Niranjan, Mahesan. On-line Q learning using connectionist systems. University of Cambridge, Department of Engineering, 1994
1994
Earlier work this paper cites.
Learning to act using real-time dynamic programming AG Barto, SJ Bradtke, SP Singh - Artificial intelligence, 1995
1995
Earlier work this paper cites.
Tsitsiklis, John N and Van Roy, Benjamin. An analysis of temporal-difference learning with function approximation. IEEE transactions on automatic control, 42(5): 674–690, 1997
1997
Earlier work this paper cites.
Sutton, Richard S and Barto, Andrew G. Reinforcement Learning: An Introduction. MIT Press Cambridge, 1998
1998
Earlier work this paper cites.
Sutton, Richard S, McAllester, David A, Singh, Satinder P, and Mansour, Yishay. Policy gradient methods for reinforcement learning with function approximation. In NIPS, volume 99, pp. 1057–1063, 1999
1999
Earlier work this paper cites.
Van Hasselt, Hado. Double Q-learning. In Lafferty, J. D., Williams, C. K. I., Shawe-Taylor, J., Zemel, R. S., and Culotta, A. (eds.), Advances in Neural Information Processing Systems 23, pp. 2613–2621. 2010
2010
Earlier work this paper cites.
2012
Cited alongside, same era.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” in Advances in Neural Information Processing Systems 25, 2012, pp. 1097– 1105
2012
Cited alongside, same era.
S. Wang and C. Manning, “Fast Dropout training,” in Proceedings of the 30th International Conference on Machine Learning. PLMR, 2013
2013
Cited alongside, same era.
Kingma, Diederik P. and Ba, Jimmy. Adam: A method for stochastic optimization. arXiv preprint arXiv: 1412.6980, 2014
2014
Cited alongside, same era.
Y. Gal and Z. Ghahramani, “Dropout as a bayesian approximation: Representing model uncertainty in deep learning,” in Proceedings of the 33rd International Conference on Machine Learning. PLMR, 2016
2016
Later among the works it cites.
2016
Later among the works it cites.
2016
Later among the works it cites.
2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2015
Cited alongside, same era.
Van Hasselt, Hado, Guez, Arthur, and Silver, David. Deep reinforcement learning with double Q-learning. arXiv preprint arXiv: 1509.06461, 2015
2015
Cited alongside, same era.
2015
Cited alongside, same era.
H. Wu and X. Gu, “Towards Dropout training for convolutional neural networks,” Neural Networks, vol. 71, no. C, pp. 1–10, 2015
2015
Cited alongside, same era.
2015
Cited alongside, same era.
Wang, Ziyu, de Freitas, Nando, and Lanctot, Marc. Dueling network architectures for deep reinforcement learning. arXiv preprint arXiv: 1511.06581, 2015
2015
Cited alongside, same era.
2017
Later among the works it cites.
Y. Gal, J. Hron, and A. Kendall, “Concrete Dropout,” in Advances in Neural Information Processing Systems 30, 2017, pp. 3581–3590
2017
Later among the works it cites.
2017
Later among the works it cites.
2019
Closest in time.
2019
Closest in time.