Fetching the paper…
Reading the bibliography…
Actor-critic (AC) methods are ubiquitous in reinforcement learning.
Asynchronous methods for deep reinforcement learning
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., and Kavukcuoglu, K. (2016) · 1937
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J. (1992) · 1992
Earlier work this paper cites.
Residual algorithms: Reinforcement learning with function approximation
Baird, L. (1995) · 1995
Earlier work this paper cites.
Actor-critic algorithms
Konda, V. R. and Tsitsiklis, J. N. (2000) · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S., McAllester, D. A., Singh, S. P., and Mansour, Y. (2000) · 2000
Earlier work this paper cites.
Simulation-based optimization of markov reward processes
Marbach, P. and Tsitsiklis, J. N. (2001) · 2001
Earlier work this paper cites.
Dual representations for dynamic programming and reinforcement learning
Wang, T., Bowling, M., and Schuurmans, D. (2007) · 2007
Earlier work this paper cites.
Natural actor-critic
Peters, J. and Schaal, S. (2008) · 2008
Earlier work this paper cites.
Convergent temporal-difference learning with arbitrary smooth function approximation
Maei, H. R., Szepesvari, C., Bhatnagar, S., Precup, D., Silver, D., and Sutton, R. S. (2009) · 2009
Earlier work this paper cites.
Fast gradient-descent methods for temporal-difference learning with linear function approximation
Sutton, R. S., Maei, H. R., Precup, D., Bhatnagar, S., Silver, D., Szepesvári, C., and Wiewiora, E. (2009) · 2009
Earlier work this paper cites.
Derivatives of logarithmic stationary distributions for policy gradient reinforcement learning
Morimura, T., Uchibe, E., Yoshimoto, J., Peters, J., and Doya, K. (2010) · 2010
Earlier work this paper cites.
Should one compute the temporal difference fix point or minimize the bellman residual? the unified oblique projection view
Scherrer, B. (2010) · 2010
Cited alongside, same era.
Incremental basis construction from temporal difference error
Sun, Y., Gomez, F., Ring, M., and Schmidhuber, J. (2011) · 2011
Cited alongside, same era.
Off-policy actor-critic
Degris, T., White, M., and Sutton, R. S. (2012) · 2012
Cited alongside, same era.
Mujoco: A physics engine for model-based control
Todorov, E., Erez, T., and Tassa, Y. (2012) · 2012
Cited alongside, same era.
Planning by prioritized sweeping with small backups
Van Seijen, H. and Sutton, R. (2013) · 2013
Cited alongside, same era.
Markov decision processes: discrete stochastic dynamic programming
Puterman, M. L. (2014) · 2014
Cited alongside, same era.
Combining policy gradient and q-learning
O’Donoghue, B., Munos, R., Kavukcuoglu, K., and Mnih, V. (2017) · 2017
Later among the works it cites.
A review on bilevel optimization: from classical to evolutionary approaches and applications
Sinha, A., Malo, P., and Deb, K. (2017) · 2017
Later among the works it cites.
Addressing function approximation error in actor-critic methods
Fujimoto, S., Hoof, H., and Meger, D. (2018) · 2018
Later among the works it cites.
Equivalence between policy gradients and soft q-learning
Schulman, J., Chen, X., and Abbeel, P. (2018) · 2018
Later among the works it cites.
Comparing direct and indirect temporal-difference methods for estimating the variance of the return
Sherstan, C., Ashley, D. R., Bennett, B., Young, K., White, A., White, M., and Sutton, R. S. (2018) · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Continuous control with deep reinforcement learning
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D. (2016) · 2016
Cited alongside, same era.
Prioritized experience replay
Schaul, T., Quan, J., Antonoglou, I., and Silver, D. (2016) · 2016
Cited alongside, same era.
High-dimensional continuous control using generalized advantage estimation
Schulman, J., Moritz, P., Levine, S., Jordan, M., and Abbeel, P. (2016) · 2016
Cited alongside, same era.
Improving stochastic policy gradients in continuous control with deep reinforcement learning using the beta distribution
Chou, P.-W., Maturana, D., and Scherer, S. (2017) · 2017
Cited alongside, same era.
Bridging the gap between value and policy based reinforcement learning
Nachum, O., Norouzi, M., Xu, K., and Schuurmans, D. (2017) · 2017
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S. (2018a)
Cited in the paper.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G. (2018) · 2018
Later among the works it cites.
A distributional code for value in dopamine-based reinforcement learning
Dabney, W., Kurth-Nelson, Z., Uchida, N., Starkweather, C. K., Hassabis, D., Munos, R., and Botvinick, M. (2020) · 2020
Later among the works it cites.
Implicit learning dynamics in stackelberg games: Equilibria characterization, convergence analysis, and empirical study
Fiez, T., Chasnov, B., and Ratliff, L. (2020) · 2020
Later among the works it cites.
An operator view of policy gradient methods
Ghosh, D., Machado, M. C., and Roux, N. L. (2020) · 2020
Later among the works it cites.
Off-policy policy gradient with stationary distribution correction
Liu, Y., Swaminathan, A., Agarwal, A., and Brunskill, E. (2020) · 2020
Later among the works it cites.
A game theoretic framework for model-based reinforcement learning
Rajeswaran, A., Mordatch, I., and Kumar, V. (2020) · 2020
Later among the works it cites.