Fetching the paper…
Reading the bibliography…
The use of target networks has been a popular and key component of recent deep Q-learning algorithms for reinforcement learning, yet little is known from the theory side.
Learning to predict by the methods of temporal differences
Sutton, R. S · 1988
Earlier work this paper cites.
Q-learning
Watkins, C. J. C. H. and Dayan, P · 1992
Earlier work this paper cites.
Residual algorithms: reinforcement learning with function approximation
Baird, L · 1995
Earlier work this paper cites.
Dynamic programming and optimal control
Bertsekas, D. P · 1995
Earlier work this paper cites.
Linear System Theory and Design
Chen, C.-T · 1995
Earlier work this paper cites.
Neuro-dynamic programming
Bertsekas, D. P. and Tsitsiklis, J. N · 1996
Earlier work this paper cites.
Linear least-squares algorithms for temporal difference learning
Bradtke, S. J. and Barto, A. G · 1996
Earlier work this paper cites.
An analysis of temporal-difference learning with function approximation
Tsitsiklis, J. N. and Van Roy, B · 1997
Earlier work this paper cites.
Convex optimization
Boyd, S. and Vandenberghe, L · 2004
Earlier work this paper cites.
A linear systems primer
Antsaklis, P. J. and Michel, A. N · 2007
Earlier work this paper cites.
Learning near-optimal policies with Bellman-residual minimization based fitted policy iteration and a single sample path
Antos, A., Szepesvári, C., and Munos, R · 2008
Earlier work this paper cites.
Projected equation methods for approximate solution of large linear systems
Bertsekas, D. P. and Yu, H · 2009
Cited alongside, same era.
Convergence results for some temporal difference methods based on least squares
Yu, H. and Bertsekas, D. P · 2009
Cited alongside, same era.
Double Q-learning
Hasselt, H. V · 2010
Cited alongside, same era.
Stochastic recursive algorithms for optimization: simultaneous perturbation methods , volume 434
Bhatnagar, S., Prasad, H. L., and Prashanth, L. A · 2012
Cited alongside, same era.
A tutorial on linear function approximators for dynamic programming and reinforcement learning
Geramifard, A., Walsh, T. J., Tellex, S., Chowdhary, G., Roy, N., How, J. P., et al · 2013
Cited alongside, same era.
Policy evaluation with temporal differences: A survey and comparison
Dann, C., Neumann, G., and Peters, J · 2014
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Later among the works it cites.
Continuous deep q-learning with model-based acceleration
Gu, S., Lillicrap, T., Sutskever, I., and Levine, S · 2016
Later among the works it cites.
Asynchronous methods for deep reinforcement learning
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., and Kavukcuoglu, K · 2016
Later among the works it cites.
Deep reinforcement learning with double Q-learning
Van Hasselt, H., Guez, A., and Silver, D · 2016
Later among the works it cites.
Dueling network architectures for deep reinforcement learning
Wang, Z., Schaul, T., Hessel, M., Hasselt, H., Lanctot, M., and Freitas, N · 2016
Later among the works it cites.
Learning from conditional distributions via dual embeddings
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Proximal reinforcement learning: A new theory of sequential decision making in primal-dual spaces
Mahadevan, S., Liu, B., Thomas, P., Dabney, W., Giguere, S., Jacek, N., Gemp, I., and Liu, J · 2014
Cited alongside, same era.
Fast lstd using stochastic approximation: Finite time analysis and application to traffic control
Prashanth, L. A., Korda, N., and Munos, R · 2014
Cited alongside, same era.
Convex optimization: Algorithms and complexity
Bubeck, S. et al · 2015
Cited alongside, same era.
Memory-based control with recurrent neural networks
Heess, N., Hunt, J. J., Lillicrap, T. P., and Silver, D · 2015
Cited alongside, same era.
Continuous control with deep reinforcement learning
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D · 2015
Cited alongside, same era.
Fast gradient-descent methods for temporal-difference learning with linear function approximation
Sutton, R. S., Maei, H. R., Precup, D., Bhatnagar, S., Silver, D., Szepesvári, C., and Wiewiora, E
Cited in the paper.
Dai, B., He, N., Pan, Y., Boots, B., and Song, L · 2017
Later among the works it cites.
A finite time analysis of temporal difference learning with linear function approximation
Bhandari, J., Russo, D., and Singal, R · 2018
Later among the works it cites.
Optimization methods for large-scale machine learning
Bottou, L., Curtis, F. E., and Nocedal, J · 2018
Later among the works it cites.
SBEED: Convergent reinforcement learning with nonlinear function approximation
Dai, B., Shaw, A., Li, L., Xiao, L., He, N., Liu, Z., Chen, J., and Song, L · 2018
Later among the works it cites.
Finite sample analyses for TD(0) with function approximation
Dalal, G., Szörényi, B., Thoppe, G., and Mannor, S · 2018
Later among the works it cites.
Finite-time error bounds for linear stochastic approximation and TD learning
Srikant, R. and Ying., L · 2019
Closest in time.