Fetching the paper…
Reading the bibliography…
Traditional exploration methods in RL require agents to perform random actions to find rewards.
Harcourt, 1924
W. Kohler, The Mentality of Apes · 1924
Earlier work this paper cites.
McGraw-Hill, 1960
D. Berlyne, Conflict, Arousal and Curiosity · 1960
Earlier work this paper cites.
R. French Trends in Cognitive Sciences
1999
Earlier work this paper cites.
A. Ng, D. Harada, and S. Russell, “Policy invariance under reward transformations: Theory and application to reward shaping,” in In Proceedings of the Sixteenth International Conference on Machine Learning
1999
Earlier work this paper cites.
S. Kakade and P. Dayan, “Dopamine: Generalization and bonuses,” Neuroscience
2000
Earlier work this paper cites.
R. Brafman and M. Tennenholtz, “R-max - a general polynomial time algorithm for near-optimal reinforcement learning.,” vol. 3, pp. 953–958, 01 2001
2001
Earlier work this paper cites.
M. Ŝpinka, R. Newberry, and M. Bekoff, “Mammalian play: Training for the unexpected,” The Quarterly Review of Biology
2001
Earlier work this paper cites.
M. Kearns and S. Singh, “Near-optimal reinforcement learning in polynomial time,” Machine Learning
2002
Earlier work this paper cites.
P. Dayan and W. Belleine, “Reward, motivation and reinforcement learning,” Neuron
2002
Earlier work this paper cites.
P.-Y. Oudeyer, F. Kaplan, and V. Hafner, “Intrinsic motivation systems for autonomous mental development,” IEEE Transactions on Evolutionary Computation
2007
Earlier work this paper cites.
J. Kolter and A. Ng, “Near-bayesian exploration in polynomial time,” International Conference on Machine Learning
2009
Earlier work this paper cites.
J. Schmidhuber, “Formal theory of creativity, fun, and intrinsic motivation (1990–2010),” IEEE Transactions on Autonomous Mental Development
2010
Earlier work this paper cites.
J. Lehman and K. Stanley, “Abandoning objectives: Evolution through the search for novelty alone,” Evolutionary Computation
2011
Earlier work this paper cites.
B. Woolley and K. Stanley, “On the deleterious effects of a priori objectives on evolution and representation,” Genetic & Evolutionary Computation Conference
2011
Cited alongside, same era.
S. Benson-Amram and K. Holekamp, “Innovative problem solving by wild spotted hyenas,” Proceedings of the Royal Society London B
2012
Cited alongside, same era.
J. Gomes and A. Christensen, “Generic behaviour similarity measures for evolutionary swarm robotics,” Genetic and Evolutionary Computation Conference
2013
Cited alongside, same era.
M. Bellemare, J. Veness, and E. Talvitie, “Skip context tree switching,” Proceedings of the 31st International Conference on Machine Learning
2014
Cited alongside, same era.
V. Mnih, K. Kavukcuoglu, D. Silver, A. Rusu, J. Veness, M. Bellemare, A. Graves, M. Riedmiller, A. Fidjeland, G. Ostrovski, S. Petersen, C. Beattie, A. Sadik, I. Antonoglou, H. King, D. Kumaran, D. Wierstra, S. Legg, and D. Hassabis, “Human-level control through deep reinforcement learning,” Nature
C. Stanton and J. Clune, “Curiosity search: Producing generalists by encouraging individuals to continually explore and acquire skills throughout their lifetime,” PloS ONE
2016
Later among the works it cites.
Accessed: 24-05-2016
“Monty python’s ministry of silly walks.” https://www.youtube.com/watch?v=iV2ViNJFZC8 · 2016
Later among the works it cites.
2017
Later among the works it cites.
H. Tang, R. Houthooft, D. Foote, A. Stooke, O. Xi Chen, Y. Duan, J. Schulman, F. DeTurck, and P. Abbeel, “#exploration: A study of count-based exploration for deep reinforcement learning,” pp. 2753–2762, Curran Associates, Inc., 2017
2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2015
Cited alongside, same era.
A. Nguyen, J. Yosinski, and J. Clune, “Innovation engines: Automated creativity and improved stochastic optimization via deep learning,” Proceedings of the Genetic and Evolutionary Computation Conference
2015
Cited alongside, same era.
K. Fragkiadaki, P. Arbelaez, P. Felsen, and J. Malik, “Learning to segment moving objects in videos,” 2015 Conference on Computer Vision and Pattern Recognition
2015
Cited alongside, same era.
A. Cully, J. Clune, D. Tarapore, and J.-B. Mouret, “Robots that can adapt like animals,” Nature
2015
Cited alongside, same era.
V. Mnih, A. Badia, M. Mirza, A. Graves, T. Harley, T. Lillicrap, D. Silver, and K. Kavukcuoglu, “Asynchronous methods for deep reinforcement learning,” in Proceedings of the 33rd International Conference on Machine Learning
2016
Cited alongside, same era.
T. Schaul, J. Quan, I. Antonoglou, and D. Silver, “Prioritized experience replay,” in International Conference on Learning Representations
2016
Cited alongside, same era.
H. van Hasselt, A. Guev, and D. Silver, “Deep reinforcement learning with double q-learning,” in Proceedings of the 30th AAAI Conference on Artificial Intelligence
2016
Cited alongside, same era.
M. Bellemare, S. Srinivasan, G. Ostrovski, T. Schaul, D. Saxton, and R. Munos, “Unifying count-based exploration and intrinsic motivation,” 30th Conference on Neural Information Processing Systems
2016
Cited alongside, same era.
2017
Later among the works it cites.
2017
Later among the works it cites.
P. Dhariwal, C. Hesse, O. Klimov, A. Nichol, M. Plappert, A. Radford, J. Schulman, S. Sidor, and Y. Wu, “Openai baselines.” https://github.com/openai/baselines , 2017
2017
Later among the works it cites.
Y. Wu, E. Mansimov, S. Liao, R. Grosse, and J. Ba, “Scalable trust-region method for deep reinforcement learning using kronecker-factored approximation,” in Advances in Neural Information Processing Systems 30
2017
Later among the works it cites.
2018
Closest in time.
2018
Closest in time.
2018
Closest in time.