Fetching the paper…
Reading the bibliography…
Designing off-policy reinforcement learning algorithms is typically a very challenging task, because a desirable iteration update often involves an expectation over an on-policy distribution.
Global convergence of policy gradient methods to (almost) locally optimal policies
Zhang, K., Koppel, A., Zhu, H., and Başar, T · 1906
Earlier work this paper cites.
Provably convergent off-policy actor-critic with function approximation
Zhang, S., Liu, B., Yao, H., and Whiteson, S · 1911
Earlier work this paper cites.
Residual algorithms: reinforcement learning with function approximation
Baird, L · 1995
Earlier work this paper cites.
Actor-critic algorithms
Konda, V. R. and Tsitsiklis, J. N · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S., McAllester, D. A., Singh, S. P., and Mansour, Y · 2000
Earlier work this paper cites.
Exploration in gradient-based reinforcement learning
Meuleau, N., Peshkin, L., and Kim, K.-E · 2001
Earlier work this paper cites.
Gendice: generalized offline estimation of stationary values
Zhang, R., Dai, B., Li, L., and Schuurmans, D · 2002
Earlier work this paper cites.
Non-asymptotic convergence analysis of two time-scale (natural) actor-critic algorithms
Xu, T., Wang, Z., and Liang, Y · 2005
Earlier work this paper cites.
Natural actor–critic algorithms
Bhatnagar, S., Sutton, R. S., Ghavamzadeh, M., and Lee, M · 2009
Earlier work this paper cites.
An actor–critic algorithm with function approximation for discounted cost constrained Markov decision processes
Bhatnagar, S · 2010
Earlier work this paper cites.
Derivatives of logarithmic stationary distributions for policy gradient reinforcement learning
Morimura, T., Uchibe, E., Yoshimoto, J., Peters, J., and Doya, K · 2010
Earlier work this paper cites.
Doubly robust policy evaluation and learning
Dudík, M., Langford, J., and Li, L · 2011
Earlier work this paper cites.
Gradient temporal-difference learning algorithms
Maei, H. R · 2011
Earlier work this paper cites.
Generalized off-policy actor-critic
Zhang, S., Boehmer, W., and Whiteson, S · 2011
Earlier work this paper cites.
Off-policy actor-critic
Degris, T., White, M., and Sutton, R. S · 2012
Earlier work this paper cites.
Doubly robust policy evaluation and optimization
Dudík, M., Erhan, D., Langford, J., Li, L., et al · 2014
Earlier work this paper cites.
Deterministic policy gradient algorithms
Silver, D., Lever, G., Heess, N., Degris, T., Wierstra, D., and Riedmiller, M · 2014
Earlier work this paper cites.
Doubly robust off-policy value evaluation for reinforcement learning
Jiang, N. and Li, L · 2016
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D · 2016
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., and Kavukcuoglu, K · 2016
Cited alongside, same era.
Combining policy gradient and q-learning
O’Donoghue, B., Munos, R., Kavukcuoglu, K., and Mnih, V · 2016
Cited alongside, same era.
Mastering the game of Go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al · 2016
Cited alongside, same era.
An emphatic approach to the problem of off-policy temporal-difference learning
Sutton, R. S., Mahmood, A. R., and White, M · 2016
Cited alongside, same era.
Sample efficient actor-critic with experience replay
Optimality and approximation with policy gradient methods in Markov decision processes
Agarwal, A., Kakade, S. M., Lee, J. D., and Mahajan, G · 2019
Later among the works it cites.
Momentum-based variance reduction in non-convex SGD
Cutkosky, A. and Orabona, F · 2019
Later among the works it cites.
Off-policy deep reinforcement learning by bootstrapping the covariate shift
Gelada, C. and Bellemare, M. G · 2019
Later among the works it cites.
Guidelines for reinforcement learning in healthcare
Gottesman, O., Johansson, F., Komorowski, M., Faisal, A., Sontag, D., Doshi-Velez, F., and Celi, L. A · 2019
Later among the works it cites.
Off-policy policy gradient with state distribution correction
Liu, Y., Swaminathan, A., Agarwal, A., and Brunskill, E · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Wang, Z., Bapst, V., Heess, N., Mnih, V., Munos, R., Kavukcuoglu, K., and de Freitas, N · 2016
Cited alongside, same era.
Reinforcement learning with deep energy-based policies
Haarnoja, T., Tang, H., Abbeel, P., and Levine, S · 2017
Cited alongside, same era.
Bridging the gap between value and policy based reinforcement learning
Nachum, O., Norouzi, M., Xu, K., and Schuurmans, D · 2017
Cited alongside, same era.
Equivalence between policy gradients and soft q-learning
Schulman, J., Chen, X., and Abbeel, P · 2017
Cited alongside, same era.
SBEED: convergent reinforcement learning with nonlinear function approximation
Dai, B., Shaw, A., Li, L., Xiao, L., He, N., Liu, Z., Chen, J., and Song, L · 2018
Cited alongside, same era.
Addressing function approximation error in actor-critic methods
Fujimoto, S., Hoof, H., and Meger, D · 2018
Cited alongside, same era.
Soft actor-critic: off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S · 2018
Cited alongside, same era.
Nachum, O., Dai, B., Kostrikov, I., Chow, Y., Li, L., and Schuurmans, D · 2019
Later among the works it cites.
Doubly robust bias reduction in infinite horizon off-policy estimation
Tang, Z., Feng, Y., Li, L., Zhou, D., and Liu, Q · 2019
Later among the works it cites.
Neural policy gradient methods: Global optimality and rates of convergence
Wang, L., Cai, Q., Yang, Z., and Wang, Z · 2019
Later among the works it cites.
Two time-scale off-policy TD learning: Non-asymptotic analysis over markovian samples
Xu, T., Zou, S., and Liang, Y · 2019
Later among the works it cites.
Single-timescale actor-critic provably finds globally optimal policy
Fu, Z., Yang, Z., and Wang, Z · 2020
Later among the works it cites.
From importance sampling to doubly robust policy gradient
Huang, J. and Jiang, N · 2020
Later among the works it cites.
Statistically efficient off-policy policy gradients
Kallus, N. and Uehara, M · 2020
Later among the works it cites.
An improved analysis of (variance-reduced) policy gradient and natural policy gradient methods
Liu, Y., Zhang, K., Basar, T., and Yin, W · 2020
Later among the works it cites.
Variance-reduced off-policy memory-efficient policy search
Lyu, D., Qi, Q., Ghavamzadeh, M., Yao, H., Yang, T., and Liu, B · 2020
Later among the works it cites.
A nonparametric off-policy policy gradient
Tosatto, S., Carvalho, J., Abdulsamad, H., and Peters, J · 2020
Later among the works it cites.
A finite time analysis of two time-scale actor critic methods
Wu, Y., Zhang, W., Xu, P., and Gu, Q · 2020
Later among the works it cites.
Non-asymptotic convergence of Adam-type reinforcement learning algorithms under Markovian sampling
Xiong, H., Xu, T., Liang, Y., and Zhang, W · 2020
Later among the works it cites.
Off-policy evaluation via the regularized lagrangian
Yang, M., Nachum, O., Dai, B., Li, L., and Schuurmans, D · 2020
Later among the works it cites.