Fetching the paper…
Reading the bibliography…
Reinforcement learning is a powerful technique to train an agent to perform a task.
Curious model-building control systems
Schmidhuber, J · 1991
Earlier work this paper cites.
Curriculum learning
Bengio, Y., Louradour, J., Collobert, R., and Weston, J · 2009
Earlier work this paper cites.
Intrinsically motivated goal exploration for active motor learning in robots: A case study
Baranes, A. and Oudeyer, P.-Y · 2010
Earlier work this paper cites.
Formal theory of creativity, fun, and intrinsic motivation (1990–2010)
Schmidhuber, J · 2010
Earlier work this paper cites.
Intrinsically motivated hierarchical skill learning in structured environments
Vigorito, C. M. and Barto, A. G · 2010
Earlier work this paper cites.
Autonomous skill acquisition on a mobile manipulator
Konidaris, G., Kuindersma, S., Grupen, R. A., and Barto, A. G · 2011
Earlier work this paper cites.
Temporal-difference competence-based intrinsic motivation (td-cb-im): A mechanism that uses the td-error as an intrinsic reinforcement for deciding which skill to learn when
Baldassarre, G. and Mirolli, M · 2012
Earlier work this paper cites.
Curriculum learning for motor skills
Karpathy, A. and Van De Panne, M · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Todorov, E., Erez, T., and Tassa, Y · 2012
Earlier work this paper cites.
A survey on policy search for robotics
Deisenroth, M. P., Neumann, G., Peters, J., et al · 2013
Earlier work this paper cites.
Data-efficient generalization of robot skills with contextual policy search
Kupcsik, A. G., Deisenroth, M. P., Peters, J., Neumann, G., and Others · 2013
Earlier work this paper cites.
Learning Stochastic Feedforward Neural Networks
Tang, Y. and Salakhutdinov, R · 2013
Earlier work this paper cites.
Multi-task policy search for robotics
Deisenroth, M. P., Englert, P., Peters, J., and Fox, D · 2014
Earlier work this paper cites.
Active contextual policy search
Fabisch, A. and Metzen, J. H · 2014
Cited alongside, same era.
Generative adversarial networks
Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y · 2014
Cited alongside, same era.
Zaremba, W. and Sutskever, I · 2014
Cited alongside, same era.
Scheduled sampling for sequence prediction with recurrent neural networks
Bengio, S., Vinyals, O., Jaitly, N., and Shazeer, N · 2015
Cited alongside, same era.
Continuous control with deep reinforcement learning
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Variational information maximizing exploration
Houthooft, R., Chen, X., Duan, Y., Schulman, J., De Turck, F., and Abbeel, P · 2016
Later among the works it cites.
End-to-end training of deep visuomotor policies
Levine, S., Finn, C., Darrell, T., and Abbeel, P · 2016
Later among the works it cites.
Strategic attentive writer for learning macro-actions
Mnih, V., Agapiou, J., Osindero, S., Graves, A., Vinyals, O., Kavukcuoglu, K., et al · 2016
Later among the works it cites.
Mastering the game of go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al · 2016
Later among the works it cites.
Value iteration networks
Tamar, A., Wu, Y., Thomas, G., Levine, S., and Abbeel, P · 2016
Later among the works it cites.
#exploration: A study of count-based exploration for deep reinforcement learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Nonparametric bayesian reward segmentation for skill discovery using inverse reinforcement learning
Ranchod, P., Rosman, B., and Konidaris, G · 2015
Cited alongside, same era.
Universal value function approximators
Schaul, T., Horgan, D., Gregor, K., and Silver, D · 2015
Cited alongside, same era.
Surprise-based intrinsic motivation for deep reinforcement learning
Achiam, J. and Sastry, S · 2016
Cited alongside, same era.
Unifying count-based exploration and intrinsic motivation
Bellemare, M. G., Srinivasan, S., Ostrovski, G., Schaul, T., Saxton, D., and Munos, R · 2016
Cited alongside, same era.
Benchmarking deep reinforcement learning for continuous control
Duan, Y., Chen, X., Houthooft, R., Schulman, J., and Abbeel, P · 2016
Cited alongside, same era.
Deep learning for reward design to improve monte carlo tree search in atari games
Guo, X., Singh, S., Lewis, R., and Lee, H · 2016
Cited alongside, same era.
Tang, H., Houthooft, R., Foote, D., Stooke, A., Chen, X., Duan, Y., Schulman, J., Turck, F. D., and Abbeel, P · 2016
Later among the works it cites.
Generating long-term trajectories using deep hierarchical networks
Zheng, S., Yue, Y., and Hobbs, J · 2016
Later among the works it cites.
Hindsight experience replay
Andrychowicz, M., Wolski, F., Ray, A., Schneider, J., Fong, R., Welinder, P., McGrew, B., Tobin, J., Abbeel, P., and Zaremba, W · 2017
Closest in time.
Wasserstein GAN
Arjovsky, M., Chintala, S., and Bottou, L · 2017
Closest in time.
Automated Curriculum Learning for Neural Networks
Graves, A., Bellemare, M. G., Menick, J., Munos, R., and Kavukcuoglu, K · 2017
Closest in time.
On the effectiveness of least squares generative adversarial networks
Mao, X., Li, Q., Xie, H., Lau, R. Y., Wang, Z., and Smolley, S. P · 2017
Closest in time.
Online Multi-Task Learning Using Biased Sampling
Sharma, S. and Ravindran, B · 2017
Closest in time.
Intrinsic Motivation and Automatic Curricula via Asymmetric Self-Play
Sukhbaatar, S., Kostrikov, I., Szlam, A., and Fergus, R · 2017
Closest in time.