Fetching the paper…
Reading the bibliography…
In importance sampling (IS)-based reinforcement learning algorithms such as Proximal Policy Optimization (PPO), IS weights are typically clipped to avoid large variance in learning.
Learning to predict by the methods of temporal differences
Sutton, R. S · 1988
Earlier work this paper cites.
Linear least-squares algorithms for temporal difference learning
Bradtke, S. J. and Barto, A. G · 1996
Earlier work this paper cites.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G · 1998
Earlier work this paper cites.
Eligibility traces for off-policy policy evaluation
Precup, D · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S., McAllester, D. A., Singh, S. P., and Mansour, Y · 2000
Earlier work this paper cites.
Importance sampling for reinforcement learning with multiple objectives
Shelton, C. R · 2001
Earlier work this paper cites.
Learning from scarce experience
Peshkin, L. and Shelton, C. R · 2002
Earlier work this paper cites.
Natural actor-critic
Peters, J., Vijayakumar, S., and Schaal, S · 2005
Earlier work this paper cites.
Real-time reinforcement learning by sequential actor–critics and experience replay
Wawrzyński, P · 2009
Earlier work this paper cites.
On a connection between importance sampling and the likelihood ratio policy gradient
Jie, T. and Abbeel, P · 2010
Earlier work this paper cites.
Box2d: A 2d physics engine for games, 2011
Catto, E · 2011
Earlier work this paper cites.
Degris, T., White, M., and Sutton, R. S · 2012
Cited alongside, same era.
Mujoco: A physics engine for model-based control
Todorov, E., Erez, T., and Tassa, Y · 2012
Cited alongside, same era.
Guided policy search
Levine, S. and Koltun, V · 2013
Cited alongside, same era.
Playing atari with deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M · 2013
Cited alongside, same era.
Continuous control with deep reinforcement learning
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D · 2015
Cited alongside, same era.
Interpolated policy gradient: Merging on-policy and off-policy gradient estimation for deep reinforcement learning
Gu, S. S., Lillicrap, T., Turner, R. E., Ghahramani, Z., Schölkopf, B., and Levine, S · 2017
Later among the works it cites.
Reinforcement learning with deep energy-based policies
Haarnoja, T., Tang, H., Abbeel, P., and Levine, S · 2017
Later among the works it cites.
Amber: Adaptive multi-batch experience replay for continuous action control
Han, S. and Sung, Y · 2017
Later among the works it cites.
Emergence of locomotion behaviours in rich environments
Heess, N., Sriram, S., Lemmon, J., Merel, J., Wayne, G., Tassa, Y., Erez, T., Wang, Z., Eslami, A., Riedmiller, M., et al · 2017
Later among the works it cites.
Sample-efficient policy optimization with stein control variate
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Cited alongside, same era.
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W · 2016
Cited alongside, same era.
Q-prop: Sample-efficient policy gradient with an off-policy critic
Gu, S., Lillicrap, T., Ghahramani, Z., Turner, R. E., and Levine, S · 2016
Cited alongside, same era.
Safe and efficient off-policy reinforcement learning
Munos, R., Stepleton, T., Harutyunyan, A., and Bellemare, M · 2016
Cited alongside, same era.
Sample efficient actor-critic with experience replay
Wang, Z., Bapst, V., Heess, N., Mnih, V., Munos, R., Kavukcuoglu, K., and de Freitas, N · 2016
Cited alongside, same era.
Openai baselines
Dhariwal, P., Hesse, C., Klimov, O., Nichol, A., Plappert, M., Radford, A., Schulman, J., Sidor, S., Wu, Y., and Zhokhov, P · 2017
Cited alongside, same era.
Off-policy policy search
Meuleau, N., Peshkin, L., Kaelbling, L. P., and Kim, K.-E
Cited in the paper.
Liu, H., Feng, Y., Mao, Y., Zhou, D., Peng, J., and Liu, Q · 2017
Later among the works it cites.
Trust-pcl: An off-policy trust region method for continuous control
Nachum, O., Norouzi, M., Xu, K., and Schuurmans, D · 2017
Later among the works it cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Later among the works it cites.
Scalable trust-region method for deep reinforcement learning using kronecker-factored approximation
Wu, Y., Mansimov, E., Grosse, R. B., Liao, S., and Ba, J · 2017
Later among the works it cites.
Impala: Scalable distributed deep-rl with importance weighted actor-learner architectures
Espeholt, L., Soyer, H., Munos, R., Simonyan, K., Mnih, V., Ward, T., Doron, Y., Firoiu, V., Harley, T., Dunning, I., et al · 2018
Later among the works it cites.
Addressing function approximation error in actor-critic methods
Fujimoto, S., van Hoof, H., and Meger, D · 2018
Later among the works it cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S · 2018
Later among the works it cites.