Fetching the paper…
Reading the bibliography…
We prove under commonly used assumptions the convergence of actor-critic reinforcement learning algorithms, which simultaneously learn a policy function, the actor, and a value function, the critic.
A stochastic approximation method
H. Robbins and S. Monro · 1951
Earlier work this paper cites.
Stochastic Approximation Methods for Constrained and Unconstrained Systems
H. J. Kushner and D. S. Clark · 1978
Earlier work this paper cites.
A dual back-propagation scheme for scalar reinforcement learning
P. W. Munro · 1987
Earlier work this paper cites.
Dynamic Error Propagation Networks
A. J. Robinson · 1989
Earlier work this paper cites.
Dynamic reinforcement driven error propagation networks with application to game playing
T. Robinson and F. Fallside · 1989
Earlier work this paper cites.
The convergence of TD( λ \lambda ) for general λ \lambda
P. Dayan · 1992
Earlier work this paper cites.
Q-Learning
C. J. C. H. Watkins and P. Dayan · 1992
Earlier work this paper cites.
Asynchronous stochastic approximation and q q -learning
J. N. Tsitsiklis · 1994
Earlier work this paper cites.
Neuro-dynamic programming
D. P. Bertsekas and J. N. Tsitsiklis · 1996
Earlier work this paper cites.
Actor-critic-type learning algorithms for Markov decision processes
V. R. Konda and V. S. Borkar · 1999
Earlier work this paper cites.
The O.D.E. method for convergence of stochastic approximation and reinforcement learning
V. S. Borkar and S. P. Meyn · 2000
Earlier work this paper cites.
Actor-critic algorithms
V. R. Konda and J. N. Tsitsiklis · 2000
Earlier work this paper cites.
Convergence results for single-step on-policy reinforcement-learning algorithms
S. Singh, T. Jaakkola, M. Littman, and C. Szepesvári · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
R. S. Sutton, D. McAllester, S. Singh, and Y. Mansour · 2000
Earlier work this paper cites.
Statistical Inference
G. Casella and R. L. Berger · 2002
Earlier work this paper cites.
On actor-critic algorithms
V. R. Konda and J. N. Tsitsiklis · 2003
Earlier work this paper cites.
Stochastic Approximation and Recursive Algorithms and Applications
H. J. Kushner and G. G. Yin · 2003
Earlier work this paper cites.
Markov Decision Processes
M. L. Puterman · 2005
Cited alongside, same era.
On the stable equilibrium points of gradient systems
P. A. Absil and K. Kurdyka · 2006
Cited alongside, same era.
On the almost sure convergence of stochastic gradient descent in non-convex problems
P. Metrikopoulos, N. Hallak, A. Kavis, and V. Cevher · 2006
Cited alongside, same era.
Reinforcement learning by backpropagation through an lstm model/critic
B. Bakker · 2007
Cited alongside, same era.
Stochastic Approximation: A Dynamical Systems Viewpoint , volume 48 of Texts and Readings in Mathematics
V. S. Borkar · 2008
Cited alongside, same era.
Convergent temporal-difference learning with arbitrary smooth function approximation
H. R. Maei, C. Szepesvári, S. Bhatnagar, D. Precup, D. Silver, and R. S. Sutton · 2009
Ergodic properties of markov processes
M. Hairer · 2018
Later among the works it cites.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2018
Later among the works it cites.
Reinforcement Learning: An Introduction
R. S. Sutton and A. G. Barto · 2018
Later among the works it cites.
Rudder: Return decomposition for delayed rewards
J. A. Arjona-Medina, M. Gillhofer, M. Widrich, T. Unterthiner, J. Brandstetter, and S. Hochreiter · 2019
Later among the works it cites.
Minmax optimization: Stable limit points of gradient descent ascent are locally optimal
C. Jin, P. Netrapalli, and M. I. Jordan · 2019
Later among the works it cites.
Depth with nonlinearity creates no bad local minima in ResNets
K. Kawaguchi and Y. Bengio · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Stochastic Recursive Algorithms for Optimization
S. Bhatnagar, H. L. Prasad, and L. A. Prashanth · 2013
Cited alongside, same era.
Playing Atari with deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. Graves, I. Antonoglou, D. Wierstra, and M. A. Riedmiller · 2013
Cited alongside, same era.
The loss surfaces of multilayer networks
A. Choromanska, M. Henaff, M. Mathieu, G. B. Arous, and Y. LeCun · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, S. Petersen, C. Beattie, A. Sadik, I. Antonoglou, H. King, D. Kumaran, D. Wierstra, S. Legg, , and D. Hassabis · 2015
Cited alongside, same era.
Trust region policy optimization
J. Schulman, S. Levine, P. Moritz, M. I. Jordan, and P. Abbeel · 2015
Cited alongside, same era.
Deep learning without poor local minima
K. Kawaguchi · 2016
Cited alongside, same era.
Effect of depth and width on local minima in deep learning
K. Kawaguchi, J. Huang, and L. P. Kaelbling · 2019
Later among the works it cites.
On gradient descent ascent for nonconvex-concave minimax problems
T. Lin, C. Jin, and M. I. Jordan · 2019
Later among the works it cites.
Neural proximal/trust region policy optimization attains globally optimal policy
B. Liu, Q. Cai, Z. Yang, and Z. Wang · 2019
Later among the works it cites.
On finding local Nash equilibria (and only local Nash equilibria) in zero-sum games
E. V. Mazumdar, M. I. Jordan, and S. S. Sastry · 2019
Later among the works it cites.
Dota 2 with large scale deep reinforcement learning
OpenAI, C. Berner, G. Brockman, B. Chan, V. Cheung, P. Debiak, C. Dennison, D. Farhi, Q. Fischer, S. Hashme, C. Hesse, R. Jozefowicz, S. Gray, C. Olsson, J. Pachocki, M. Petrov, H. P. deOliveira Pinto, J. Raiman, T. Salimans, J. Schlatter, J. Schneider, S. Sidor, I. Sutskever, J. Tang, F. Wolski, and S. Zhang · 2019
Later among the works it cites.
Grandmaster level in StarCraft II using multi-agent reinforcement learning
O. Vinyals, I. Babuschkin, W. Czarnecki, M. Mathieu, A. Dudzik, J. Chung, D. Choi, R. Powell, T. Ewalds, P. Georgiev, J. Oh, D. Horgan, M. Kroiss, I. Danihelka, A. Huang, L. Sifre, T. Cai, J. Agapiou, M. Jaderberg, and D. Silver · 2019
Later among the works it cites.
Two time-scale off-policy td learning: Non-asymptotic analysis over Markovian samples
T. Xu, S. Zou, and Y Liang · 2019
Later among the works it cites.
Provably global convergence of actor-critic: A case for linear quadratic regulator with ergodic cost
Z. Yang, Y. Chen, M. Hong, and Z. Wang · 2019
Later among the works it cites.
A theoretical analysis of deep q q -learning
J. Fan, Z. Wang, Y. Xie, and Z. Yang · 2020
Closest in time.
Align-rudder: Learning from few demonstrations by reward redistribution
V. P. Patil, M. Hofmarcher, M. C. Dinu, M. Dorfer, P. Blies, J. Brandstetter, J. A. Arjona-Medina, and S. Hochreiter · 2020
Closest in time.