Fetching the paper…
Reading the bibliography…
The research on deep reinforcement learning which estimates Q-value by deep learning has been attracted the interest of researchers recently.
Richard Bellman. On the Theory of Dynamic Programming. In PNAS, 38(8):716-719, 1952
1952
Earlier work this paper cites.
Bentley, Jon Louis. Multidimensional binary search trees used for associative searching. Commun. ACM, 18(9): 509-517, 1975
1975
Earlier work this paper cites.
Long-Ji Lin. Self-Improving Reactive Agents Based on Reinforcement Learning, Planning and Teaching. Machine Learning, 8(3-4):293-321,1992
1992
Earlier work this paper cites.
Christopher JCH Watkins and Peter Dayan. Q-Learning. Machine Learning, 8(3-4):279-292, 1992
1992
Earlier work this paper cites.
Jing Peng and Ronald Williams. Incremental multi-step Q-learning. Machine Learning, 22:283- 290, 1996
1996
Earlier work this paper cites.
Hado van Hasselt. Double Q-Learning. In NIPS, 2010
2010
Earlier work this paper cites.
Alex Krizhevsky, Ilya Sutskever, Geoffrey E,Hinton. ImageNet classification with deep convolutional neural networks. In ANIPS, 2012
2012
Cited alongside, same era.
Mnih, Volodymyr, Kavukcuoglu, Koray, Silver, David, Rusu, Andrei A, Veness, Joel, Bellemare, Marc G, Graves, Alex, Riedmiller, Martin, Fidjeland, Andreas K, Ostrovski, Georg, et al. Human-level control through deep reinforcement learning. In Nature, 518(7540):529-533, 2015
2015
Cited alongside, same era.
Schaul, Tom Schaul, John Quan, Ioannis Antonoglou, and David Silver. Prioritized Experience Replay. In ICLR, 2016
2016
Cited alongside, same era.
Hado van Hasselt, Arthur Guez, and David Silver. Deep Reinforcement Learning with Double Q-Learning. In AAAI, 2016
2016
Cited alongside, same era.
2016
Later among the works it cites.
Hado van Hasselt, Arthur Guez, Matteo Hessel, Volodymyr Mnih, David Silver. Learning values across many orders of magnitudes. In NIPS, 2016
2016
Later among the works it cites.
2017
Later among the works it cites.
2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2016
Cited alongside, same era.
Remi Munos, Tom Stepleton, Anna Harutyunyan, and Marc G Bellemare. Safe and Efficient Off-Policy Reinforcement Learning. In NIPS, 2016
2016
Cited alongside, same era.
Shoya Matsumori, Takuma Seno, Toshiki Kikuchi, Yusuke Takimoto, Masahiko Osawa, Michita Imai. Embedding Cognitive Map in Neural Episodic Control. In SIG-AGI, 2017
2017
Later among the works it cites.