Fetching the paper…
Reading the bibliography…
The purpose of this technical report is two-fold.
Mujoco: A physics engine for model-based control
Todorov, E., Erez, T., and Tassa, Y. (2012) · 2012
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. and Ba, J. (2014) · 2014
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D. (2015) · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al. (2015) · 2015
Earlier work this paper cites.
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W. (2016) · 2016
Earlier work this paper cites.
Learning to play in a day: Faster deep reinforcement learning by optimality tightening
He, F. S., Liu, Y., Schwing, A. G., and Peng, J. (2016) · 2016
Earlier work this paper cites.
Safe and efficient off-policy reinforcement learning
Munos, R., Stepleton, T., Harutyunyan, A., and Bellemare, M. (2016) · 2016
Cited alongside, same era.
Hindsight experience replay
Andrychowicz, M., Wolski, F., Ray, A., Schneider, J., Fong, R., Welinder, P., McGrew, B., Tobin, J., Abbeel, O. P., and Zaremba, W. (2017) · 2017
Cited alongside, same era.
A distributional perspective on reinforcement learning
Bellemare, M. G., Dabney, W., and Munos, R. (2017) · 2017
Cited alongside, same era.
OpenAI Baselines
Dhariwal, P., Hesse, C., Klimov, O., Nichol, A., Plappert, M., Radford, A., Schulman, J., Sidor, S., and Wu, Y. (2017) · 2017
Cited alongside, same era.
Reverse curriculum generation for reinforcement learning
Florensa, C., Held, D., Wulfmeier, M., and Abbeel, P. (2017) · 2017
Cited alongside, same era.
Gu, S., Lillicrap, T., Ghahramani, Z., Turner, R. E., Schölkopf, B., and Levine, S. (2017) · 2017
Later among the works it cites.
Levy, A., Platt, R., and Saenko, K. (2017) · 2017
Later among the works it cites.
Parameter space noise for exploration
Plappert, M., Houthooft, R., Dhariwal, P., Sidor, S., Chen, R. Y., Chen, X., Asfour, T., Abbeel, P., and Andrychowicz, M. (2017) · 2017
Later among the works it cites.
Rauber, P., Mutz, F., and Schmidhuber, J. (2017) · 2017
Later among the works it cites.
Temporal difference models: Model-free deep rl for model-based control
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Schaul, T., Horgan, D., Gregor, K., and Silver, D. (2015a)
Cited in the paper.
Schaul, T., Quan, J., Antonoglou, I., and Silver, D. (2015b)
Cited in the paper.
Equivalence between policy gradients and soft q-learning
Schulman, J., Abbeel, P., and Chen, X. (2017a)
Cited in the paper.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O. (2017b)
Cited in the paper.
Pong, V., Gu, S., Dalal, M., and Levine, S. (2018) · 2018
Closest in time.