Understand
We want to make progress toward artificial general intelligence, namely general-purpose agents that autonomously learn how to competently act in complex environments.
- The purpose of this report is to sketch a research outline, share some of the most important open issues we are facing, and stimulate further discussion in the community.
- The content is based on some of our discussions during a week-long workshop held in Barbados in February 2018.
Built on
A possibility for implementing curiosity and boredom in model-building neural controllers
J. Schmidhuber · 1991
Earlier work this paper cites.
Continual learning in reinforcement environments
M. B. Ring · 1994
Earlier work this paper cites.
Predictive representations of state
M. L. Littman and R. S. Sutton · 2002
Earlier work this paper cites.
Memory retention–the synaptic stability versus plasticity dilemma
W. C. Abraham and A. Robins · 2005
Earlier work this paper cites.
Horde: A scalable real-time architecture for learning knowledge from unsupervised sensorimotor interaction
R. S. Sutton, J. Modayil, M. Delp, T. Degris, P. M. Pilarski, A. White, and D. Precup · 2011
Earlier work this paper cites.
Scaling-up knowledge for a cognizant robot
T. Degris and J. Modayil · 2012
Earlier work this paper cites.
Similar
Ella: An efficient lifelong learning algorithm
P. Ruvolo and E. Eaton · 2013
Cited alongside, same era.
Better Generalization with Forecasts
T. Schaul and M. B. Ring · 2013
Cited alongside, same era.
Variational information maximisation for intrinsically motivated reinforcement learning
S. Mohamed and D. J. Rezende · 2015
Cited alongside, same era.
Universal value function approximators
T. Schaul, D. Horgan, K. Gregor, and D. Silver · 2015
Cited alongside, same era.
Reinforcement learning with unsupervised auxiliary tasks
M. Jaderberg, V. Mnih, W. M. Czarnecki, T. Schaul, J. Z. Leibo, D. Silver, and K. Kavukcuoglu · 2016
Cited alongside, same era.
Safe and efficient off-policy reinforcement learning
R. Munos, T. Stepleton, A. Harutyunyan, and M. Bellemare · 2016
Cited alongside, same era.
Then
A distributional perspective on reinforcement learning
M. G. Bellemare, W. Dabney, and R. Munos · 2017
Later among the works it cites.
Multi-objective decision making
D. M. Roijers and S. Whiteson · 2017
Later among the works it cites.
The predictron: End-to-end learning and planning
D. Silver, H. van Hasselt, M. Hessel, T. Schaul, A. Guez, T. Harley, G. Dulac-Arnold, D. Reichert, N. Rabinowitz, A. Barreto, and T. Degris · 2017
Later among the works it cites.
Learning to search with mctsnets
A. Guez, T. Weber, I. Antonoglou, K. Simonyan, O. Vinyals, D. Wierstra, R. Munos, and D. Silver · 2018
Closest in time.
Reinforcement learning: An introduction
R. S. Sutton and A. G. Barto · 2018
Closest in time.
Beyond the bibliography
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…