Fetching the paper…
Reading the bibliography…
On-policy reinforcement learning (RL) algorithms have high sample complexity while off-policy algorithms are difficult to tune.
Asynchronous methods for deep reinforcement learning
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., and Kavukcuoglu, K. (2016) · 1937
Earlier work this paper cites.
A note on importance sampling using standardized weights
Kong, A. (1992) · 1992
Earlier work this paper cites.
Self-improving reactive agents based on reinforcement learning, planning and teaching
Lin, L.-J. (1992) · 1992
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J. (1992) · 1992
Earlier work this paper cites.
Actor-critic algorithms
Konda, V. R. and Tsitsiklis, J. N. (2000) · 2000
Earlier work this paper cites.
Infinite-horizon policy-gradient estimation
Baxter, J. and Bartlett, P. L. (2001) · 2001
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Kakade, S. and Langford, J. (2002) · 2002
Earlier work this paper cites.
Learning from scarce experience
Peshkin, L. and Shelton, C. R. (2002) · 2002
Earlier work this paper cites.
Doubly robust estimation in missing data and causal inference models
Bang, H. and Robins, J. M. (2005) · 2005
Earlier work this paper cites.
A basic formula for online policy gradient algorithms
Cao, X.-R. (2005) · 2005
Earlier work this paper cites.
Demystifying double robustness: A comparison of alternative strategies for estimating a population mean from incomplete data
Kang, J. D., Schafer, J. L., et al. (2007) · 2007
Earlier work this paper cites.
Dataset shift in machine learning
Quionero-Candela, J., Sugiyama, M., Schwaighofer, A., and Lawrence, N. D. (2009) · 2009
Earlier work this paper cites.
On a connection between importance sampling and the likelihood ratio policy gradient
Jie, T. and Abbeel, P. (2010) · 2010
Earlier work this paper cites.
Doubly robust policy evaluation and learning
Dudík, M., Langford, J., and Li, L. (2011) · 2011
Earlier work this paper cites.
Degris, T., White, M., and Sutton, R. S. (2012) · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Todorov, E., Erez, T., and Tassa, Y. (2012) · 2012
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M. (2013) · 2013
Cited alongside, same era.
Guided policy search
Levine, S. and Koltun, V. (2013) · 2013
Cited alongside, same era.
Playing atari with deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M. (2013) · 2013
Cited alongside, same era.
A probability path
Resnick, S. I. (2013) · 2013
Cited alongside, same era.
Monte Carlo statistical methods
Robert, C. and Casella, G. (2013) · 2013
Cited alongside, same era.
Sequential Monte Carlo methods in practice
Smith, A. (2013) · 2013
Cited alongside, same era.
Safe and efficient off-policy reinforcement learning
Munos, R., Stepleton, T., Harutyunyan, A., and Bellemare, M. (2016) · 2016
Later among the works it cites.
Deep reinforcement learning with double q-learning
Van Hasselt, H., Guez, A., and Silver, D. (2016) · 2016
Later among the works it cites.
Sample efficient actor-critic with experience replay
Wang, Z., Bapst, V., Heess, N., Mnih, V., Munos, R., Kavukcuoglu, K., and de Freitas, N. (2016) · 2016
Later among the works it cites.
A distributional perspective on reinforcement learning
Bellemare, M. G., Dabney, W., and Munos, R. (2017) · 2017
Later among the works it cites.
Interpolated policy gradient: Merging on-policy and off-policy gradient estimation for deep reinforcement learning
Gu, S. S., Lillicrap, T., Turner, R. E., Ghahramani, Z., Schölkopf, B., and Levine, S. (2017) · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Backward q-learning: The combination of sarsa algorithm and q-learning
Wang, Y.-H., Li, T.-H. S., and Lin, C.-J. (2013) · 2013
Cited alongside, same era.
Deterministic policy gradient algorithms
Silver, D., Lever, G., Heess, N., Degris, T., Wierstra, D., and Riedmiller, M. (2014) · 2014
Cited alongside, same era.
Bias in natural actor-critic algorithms
Thomas, P. (2014) · 2014
Cited alongside, same era.
Policy gradient methods for off-policy control
Lehnert, L. and Precup, D. (2015) · 2015
Cited alongside, same era.
Continuous control with deep reinforcement learning
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D. (2015) · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al. (2015) · 2015
Cited alongside, same era.
Plato: Policy learning using adaptive trajectory optimization
Kahn, G., Zhang, T., Levine, S., and Abbeel, P. (2017) · 2017
Later among the works it cites.
Action-depedent control variates for policy optimization via stein’s identity
Liu, H., Feng, Y., Mao, Y., Zhou, D., Peng, J., and Liu, Q. (2017) · 2017
Later among the works it cites.
Machado, M. C., Bellemare, M. G., Talvitie, E., Veness, J., Hausknecht, M. J., and Bowling, M. (2017) · 2017
Later among the works it cites.
Bridging the gap between value and policy based reinforcement learning
Nachum, O., Norouzi, M., Xu, K., and Schuurmans, D. (2017) · 2017
Later among the works it cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O. (2017) · 2017
Later among the works it cites.
Addressing function approximation error in actor-critic methods
Fujimoto, S., van Hoof, H., and Meger, D. (2018) · 2018
Later among the works it cites.
Deep reinforcement learning that matters
Henderson, P., Islam, R., Bachman, P., Pineau, J., Precup, D., and Meger, D. (2018) · 2018
Later among the works it cites.
Are deep policy gradient algorithms truly policy gradient algorithms?
Ilyas, A., Engstrom, L., Santurkar, S., Tsipras, D., Janoos, F., Rudolph, L., and Madry, A. (2018) · 2018
Later among the works it cites.
Oh, J., Guo, Y., Singh, S., and Lee, H. (2018) · 2018
Later among the works it cites.
The mirage of action-dependent baselines in reinforcement learning
Tucker, G., Bhupatiraju, S., Gu, S., Turner, R. E., Ghahramani, Z., and Levine, S. (2018) · 2018
Later among the works it cites.