Fetching the paper…
Reading the bibliography…
The use of target networks is a common practice in deep reinforcement learning for stabilizing the training; however, theoretical understanding of this technique is still limited.
Q-learning
Christopher J. C. H. Watkins and Peter Dayan · 1992
Earlier work this paper cites.
Reinforcement learning: An introduction
Richard S. Sutton and Andrew G. Barto · 1998
Earlier work this paper cites.
Finite-sample convergence rates for q-learning and indirect algorithms
Michael J Kearns and Satinder P Singh · 1999
Earlier work this paper cites.
Learning rates for Q-learning
Eyal Even-Dar and Yishay Mansour · 2003
Earlier work this paper cites.
Tree-based batch mode reinforcement learning
Damien Ernst, Pierre Geurts, and Louis Wehenkel · 2005
Earlier work this paper cites.
PAC model-free reinforcement learning
Alexander L Strehl, Lihong Li, Eric Wiewiora, John Langford, and Michael L Littman · 2006
Earlier work this paper cites.
Reinforcement learning in finite MDPs: PAC analysis
Alexander L Strehl, Lihong Li, and Michael L Littman · 2009
Earlier work this paper cites.
Double Q-learning
Hado V Hasselt · 2010
Earlier work this paper cites.
Ergodic mirror descent
John C Duchi, Alekh Agarwal, Mikael Johansson, and Michael I Jordan · 2012
Cited alongside, same era.
Minimax PAC bounds on the sample complexity of reinforcement learning with a generative model
Mohammad Gheshlaghi Azar, Rémi Munos, and Hilbert J Kappen · 2013
Cited alongside, same era.
Accelerating stochastic gradient descent using predictive variance reduction
Rie Johnson and Tong Zhang · 2013
Cited alongside, same era.
Adam: a method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Cited alongside, same era.
Markov decision processes: Discrete stochastic dynamic programming
Martin L. Puterman · 2014
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K. Fidjeland, Georg Ostrovski, et al · 2015
Optimization methods for large-scale machine learning
Léon Bottou, Frank E Curtis, and Jorge Nocedal · 2018
Later among the works it cites.
Finite sample analyses for TD(0) with function approximation
Gal Dalal, Balázs Szörényi, Gugan Thoppe, and Shie Mannor · 2018
Later among the works it cites.
Stochastic primal-dual Q-learning
Donghwan Lee and Niao He · 2018
Later among the works it cites.
Near-optimal time and sample complexities for solving markov decision processes with a generative model
Aaron Sidford, Mengdi Wang, Xian Wu, Lin Yang, and Yinyu Ye · 2018
Later among the works it cites.
On markov chain gradient descent
Tao Sun, Yuejiao Sun, and Wotao Yin · 2018
Later among the works it cites.
Target-based temporal-difference learning
Donghwan Lee and Niao He · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Katyusha: the first direct acceleration of stochastic gradient methods
Zeyuan Allen-Zhu · 2017
Cited alongside, same era.
Zap Q-learning
Adithya M Devraj and Sean Meyn · 2017
Cited alongside, same era.
SAGA: A fast incremental gradient method with support for non-strongly convex composite objectives
Aaron Defazio, Francis Bach, and Simon Lacoste-Julien
Cited in the paper.
Finito: A faster, permutable incremental gradient method for big data problems
Aaron Defazio, Justin Domke, et al
Cited in the paper.
Later among the works it cites.
A theoretical analysis of deep Q-learning
Zhuora Yang, Yuchen Xie, and Zhaoran Wang · 2019
Later among the works it cites.