Fetching the paper…
Reading the bibliography…
In reinforcement learning, it is common to let an agent interact for a fixed amount of time with its environment before resetting it and repeating the process in a series of episodes.
Learning to predict by the methods of temporal differences
Sutton, R. S · 1988
Earlier work this paper cites.
A survey of algorithmic methods for partially observed Markov decision processes
Lovejoy, W. S · 1991
Earlier work this paper cites.
Learning to perceive and act by trial and error
Whitehead, S. D. and Ballard, D. H · 1991
Earlier work this paper cites.
Self-improving reactive agents based on reinforcement learning, planning and teaching
Lin, L.-J · 1992
Earlier work this paper cites.
Q-learning
Watkins, C. J. and Dayan, P · 1992
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J · 1992
Earlier work this paper cites.
Dynamic programming and optimal control
Bertsekas, D. P · 1995
Earlier work this paper cites.
Neuro-Dynamic Programming
Bertsekas, D. P. and Tsitsiklis, J · 1996
Earlier work this paper cites.
Reinforcement learning: a survey
Kaelbling, L. P., Littman, M. L., and Moore, A. W · 1996
Earlier work this paper cites.
Reinforcement learning with time
Harada, D · 1997
Earlier work this paper cites.
Reinforcement Learning: an Introduction
Sutton, R. S. and Barto, A. G · 1998
Earlier work this paper cites.
Policy invariance under reward transformations: theory and application to reward shaping
Ng, A. Y., Harada, D., and Russell, S · 1999
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S., McAllester, D. A., Singh, S. P., and Mansour, Y · 2000
Cited alongside, same era.
Recurrent policy gradients
Wierstra, D., Förster, A., Peters, J., and Schmidhuber, J · 2009
Cited alongside, same era.
Algorithms for Reinforcement Learning
Szepesvari, C · 2010
Cited alongside, same era.
MuJoCo: a physics engine for model-based control
Todorov, E., Erez, T., and Tassa, Y · 2012
Cited alongside, same era.
The Arcade Learning Environment: an evaluation platform for general agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M · 2013
Cited alongside, same era.
Markov decision processes: discrete stochastic dynamic programming
Puterman, M. L · 2014
Cited alongside, same era.
Prioritized experience replay
Schaul, T., Quan, J., Antonoglou, I., and Silver, D · 2016
Later among the works it cites.
High-dimensional continuous control using generalized advantage estimation
Schulman, J., Moritz, P., Levine, S., Jordan, M., and Abbeel, P · 2016
Later among the works it cites.
A brief survey of deep reinforcement learning
Arulkumaran, K., Deisenroth, M. P., Brundage, M., and Bharath, A. A · 2017
Closest in time.
Reverse curriculum generation for reinforcement learning
Florensa, C., Held, D., Wulfmeier, M., and Abbeel, P · 2017
Closest in time.
Emergence of locomotion behaviours in rich environments
Heess, N., Sriram, S., Lemmon, J., Merel, J., Wayne, G., Tassa, Y., Erez, T., Wang, Z., Eslami, A., Riedmiller, M., and Silver, D · 2017
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deep learning
LeCun, Y., Bengio, Y., and Hinton, G · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., Petersen, S., Beattie, C., Sadik, A., Antonoglou, I., King, H., Kumaran, D., Wierstra, D., Legg, S., and Hassabis, D · 2015
Cited alongside, same era.
Deep learning in neural networks: an overview
Schmidhuber, J · 2015
Cited alongside, same era.
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W · 2016
Cited alongside, same era.
Benchmarking deep reinforcement learning for continuous control
Duan, Y., Chen, X., Houthooft, R., Schulman, J., and Abbeel, P · 2016
Cited alongside, same era.
Deep Learning
Goodfellow, I., Bengio, Y., and Courville, A · 2016
Cited alongside, same era.
Deep reinforcement learning that matters
Henderson, P., Islam, R., Bachman, P., Pineau, J., Precup, D., and Meger, D · 2017
Closest in time.
OpenAI Baselines
Hesse, C., Plappert, M., Radford, A., Schulman, J., Sidor, S., and Wu, Y · 2017
Closest in time.
Machado, M. C., Bellemare, M. G., Talvitie, E., Veness, J., Hausknecht, M., and Bowling, M · 2017
Closest in time.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Closest in time.
Unifying task specification in reinforcement learning
White, M · 2017
Closest in time.
A deeper look at experience replay
Zhang, S. and Sutton, R. S · 2017
Closest in time.