Fetching the paper…
Reading the bibliography…
Continually solving new, unsolved tasks is the key to learning diverse behaviors.
Teaching machines
B. F. Skinner · 1958
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
R. J. Williams · 1992
Earlier work this paper cites.
Learning and development in neural networks: The importance of starting small
J. L. Elman · 1993
Earlier work this paper cites.
Reinforcement learning: A survey
L. P. Kaelbling, M. L. Littman, and A. W. Moore · 1996
Earlier work this paper cites.
The acquisition of skilled motor performance: fast and slow experience-driven changes in primary motor cortex
A. Karni, G. Meyer, C. Rey-Hipolito, P. Jezzard, M. M. Adams, R. Turner, and L. G. Ungerleider · 1998
Earlier work this paper cites.
Modulation of the palmar grasp behavior in neonates according to texture property
M. Molina and F. Jouen · 1998
Earlier work this paper cites.
Introduction to reinforcement learning , volume 2
R. S. Sutton, A. G. Barto, et al · 1998
Earlier work this paper cites.
Off-policy temporal-difference learning with function approximation
D. Precup, R. S. Sutton, and S. Dasgupta · 2001
Earlier work this paper cites.
A natural policy gradient
S. M. Kakade · 2002
Earlier work this paper cites.
Abalearn: Efficient self-play learning of the game abalone
P. Campos and T. Langlois · 2003
Earlier work this paper cites.
A day of great illumination: Bf skinner’s discovery of shaping
G. B. Peterson · 2004
Earlier work this paper cites.
Bayesian neural networks for internet traffic classification
T. Auld, A. W. Moore, and S. F. Gull · 2007
Earlier work this paper cites.
R-iac: Robust intrinsically motivated exploration and active learning
A. Baranes and P.-Y. Oudeyer · 2009
Earlier work this paper cites.
Curriculum learning
Y. Bengio, J. Louradour, R. Collobert, and J. Weston · 2009
Earlier work this paper cites.
Flexible shaping: How learning in small steps helps
K. A. Krueger and P. Dayan · 2009
Cited alongside, same era.
Self-paced learning for latent variable models
M. P. Kumar, B. Packer, and D. Koller · 2010
Cited alongside, same era.
Active learning of inverse models with intrinsically motivated goal exploration in robots
A. Baranes and P.-Y. Oudeyer · 2013
Cited alongside, same era.
Self-organization of early vocal development in infants and machines: the role of intrinsic motivation
C. Moulin-Frier, S. M. Nguyen, and P.-Y. Oudeyer · 2014
Cited alongside, same era.
Deterministic policy gradient algorithms
D. Silver, G. Lever, N. Heess, T. Degris, D. Wierstra, and M. Riedmiller · 2014
Cited alongside, same era.
Robust adversarial reinforcement learning
L. Pinto, J. Davidson, R. Sukthankar, and A. Gupta · 2017
Later among the works it cites.
Mastering the game of go without human knowledge
D. Silver, J. Schrittwieser, K. Simonyan, I. Antonoglou, A. Huang, A. Guez, T. Hubert, L. Baker, M. Lai, A. Bolton, et al · 2017
Later among the works it cites.
Intrinsic motivation and automatic curricula via asymmetric self-play
S. Sukhbaatar, Z. Lin, I. Kostrikov, G. Synnaeve, A. Szlam, and R. Fergus · 2017
Later among the works it cites.
Soft actor-critic algorithms and applications
T. Haarnoja, A. Zhou, K. Hartikainen, G. Tucker, S. Ha, J. Tan, V. Kumar, H. Zhu, A. Gupta, P. Abbeel, et al · 2018
Later among the works it cites.
Deep drone racing: Learning agile flight in dynamic environments
E. Kaufmann, A. Loquercio, R. Ranftl, A. Dosovitskiy, V. Koltun, and D. Scaramuzza · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra · 2015
Cited alongside, same era.
Universal value function approximators
T. Schaul, D. Horgan, K. Gregor, and D. Silver · 2015
Cited alongside, same era.
Trust region policy optimization
J. Schulman, S. Levine, P. Abbeel, M. I. Jordan, and P. Moritz · 2015
Cited alongside, same era.
Dropout as a bayesian approximation: Representing model uncertainty in deep learning
Y. Gal and Z. Ghahramani · 2016
Cited alongside, same era.
Vime: Variational information maximizing exploration
R. Houthooft, X. Chen, Y. Duan, J. Schulman, F. De Turck, and P. Abbeel · 2016
Cited alongside, same era.
Hindsight experience replay
M. Andrychowicz, F. Wolski, A. Ray, J. Schneider, R. Fong, P. Welinder, B. McGrew, J. Tobin, P. Abbeel, and W. Zaremba · 2017
Cited alongside, same era.
Emergent complexity via multi-agent competition
T. Bansal, J. Pachocki, S. Sidor, I. Sutskever, and I. Mordatch · 2017
Cited alongside, same era.
Later among the works it cites.
Cassl: Curriculum accelerated self-supervised learning
A. Murali, L. Pinto, D. Gandhi, and A. Gupta · 2018
Later among the works it cites.
Multi-goal reinforcement learning: Challenging robotics environments and request for research
M. Plappert, M. Andrychowicz, A. Ray, B. McGrew, B. Baker, G. Powell, J. Schneider, J. Tobin, M. Chociej, P. Welinder, et al · 2018
Later among the works it cites.
Solving rubik’s cube with a robot hand
I. Akkaya, M. Andrychowicz, M. Chociej, M. Litwin, B. McGrew, A. Petron, A. Paino, M. Plappert, G. Powell, R. Ribas, et al · 2019
Later among the works it cites.
Emergent tool use from multi-agent autocurricula
B. Baker, I. Kanitscheider, T. Markov, Y. Wu, G. Powell, B. McGrew, and I. Mordatch · 2019
Later among the works it cites.
Dota 2 with large scale deep reinforcement learning
C. Berner, G. Brockman, B. Chan, V. Cheung, P. Dębiak, C. Dennison, D. Farhi, Q. Fischer, S. Hashme, C. Hesse, et al · 2019
Later among the works it cites.
Self-supervised exploration via disagreement
D. Pathak, D. Gandhi, and A. Gupta · 2019
Later among the works it cites.
Teacher algorithms for curriculum learning of deep rl in continuously parameterized environments
R. Portelas, C. Colas, K. Hofmann, and P.-Y. Oudeyer · 2019
Later among the works it cites.
Learning dexterous in-hand manipulation
O. M. Andrychowicz, B. Baker, M. Chociej, R. Jozefowicz, B. McGrew, J. Pachocki, A. Petron, M. Plappert, G. Powell, A. Ray, et al · 2020
Closest in time.