Fetching the paper…
Reading the bibliography…
We hypothesize that curiosity is a mechanism found by evolution that encourages meaningful exploration early in an agent's life in order to expose it to experiences that enable it to obtain high rewards over the course of its lifetime.
Evolutionary principles in self-referential learning, or on learning how to learn: the meta-meta-… hook
Jürgen Schmidhuber · 1987
Earlier work this paper cites.
On the search for new learning rules for anns
Samy Bengio, Yoshua Bengio, and Jocelyn Cloutier · 1995
Earlier work this paper cites.
Learning to learn
Sebastian Thrun and Lorien Pratt · 1998
Earlier work this paper cites.
Evolving neural networks through augmenting topologies
Kenneth O Stanley and Risto Miikkulainen · 2002
Earlier work this paper cites.
Intrinsic motivation systems for autonomous mental development
Pierre-Yves Oudeyer, Frdric Kaplan, and Verena V Hafner · 2007
Earlier work this paper cites.
Exploiting open-endedness to solve problems through the search for novelty
Joel Lehman and Kenneth O Stanley · 2008
Earlier work this paper cites.
Driven by compression progress: A simple principle explains essential aspects of subjective beauty, novelty, surprise, interestingness, attention, curiosity, creativity, art, science, music, jokes
Jürgen Schmidhuber · 2008
Earlier work this paper cites.
Satenstein: Automatically building local search sat solvers from components
Ashiqur R KhudaBukhsh, Lin Xu, Holger H Hoos, and Kevin Leyton-Brown · 2009
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Earlier work this paper cites.
Almost optimal exploration in multi-armed bandits
Zohar Karnin, Tomer Koren, and Oren Somekh · 2013
Earlier work this paper cites.
First experiments with powerplay
Rupesh Kumar Srivastava, Bas R Steunebrink, and Jürgen Schmidhuber · 2013
Earlier work this paper cites.
Probably Approximately Correct: NatureÕs Algorithms for Learning and Prospering in a Complex World
Leslie Valiant · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Efficient and robust automated machine learning
Matthias Feurer, Aaron Klein, Katharina Eggensperger, Jost Springenberg, Manuel Blum, and Frank Hutter · 2015
Earlier work this paper cites.
Actor-mimic: Deep multitask and transfer reinforcement learning
Emilio Parisotto, Jimmy Lei Ba, and Ruslan Salakhutdinov · 2015
Earlier work this paper cites.
Neural programmer-interpreters
Scott Reed and Nando De Freitas · 2015
Earlier work this paper cites.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Earlier work this paper cites.
Rl2: Fast reinforcement learning via slow reinforcement learning
Yan Duan, John Schulman, Xi Chen, Peter L Bartlett, Ilya Sutskever, and Pieter Abbeel · 2016
Earlier work this paper cites.
Modular active curiosity-driven discovery of tool use
Sébastien Forestier and Pierre-Yves Oudeyer · 2016
Earlier work this paper cites.
Non-stochastic best arm identification and hyperparameter optimization
Kevin Jamieson and Ameet Talwalkar · 2016
Earlier work this paper cites.
Hierarchical deep reinforcement learning: Integrating temporal abstraction and intrinsic motivation
Tejas D Kulkarni, Karthik Narasimhan, Ardavan Saeedi, and Josh Tenenbaum · 2016
Earlier work this paper cites.
Hyperband: A novel bandit-based approach to hyperparameter optimization
Lisha Li, Kevin Jamieson, Giulia DeSalvo, Afshin Rostamizadeh, and Ameet Talwalkar · 2016
Earlier work this paper cites.
Towards automatically-tuned neural networks
Hector Mendoza, Aaron Klein, Matthias Feurer, Jost Tobias Springenberg, and Frank Hutter · 2016
Earlier work this paper cites.
Andrei A Rusu, Neil C Rabinowitz, Guillaume Desjardins, Hubert Soyer, James Kirkpatrick, Koray Kavukcuoglu, Razvan Pascanu, and Raia Hadsell · 2016
Earlier work this paper cites.
Neural architecture search with reinforcement learning
Barret Zoph and Quoc V Le · 2016
Earlier work this paper cites.
Neural optimizer search with reinforcement learning
Irwan Bello, Barret Zoph, Vijay Vasudevan, and Quoc V Le · 2017
Cited alongside, same era.
Pathnet: Evolution channels gradient descent in super neural networks
Chrisantha Fernando, Dylan Banarse, Charles Blundell, Yori Zwols, David Ha, Andrei A Rusu, Alexander Pritzel, and Daan Wierstra · 2017
Cited alongside, same era.
Model-agnostic meta-learning for fast adaptation of deep networks
Chelsea Finn, Pieter Abbeel, and Sergey Levine · 2017
Cited alongside, same era.
Stochastic neural networks for hierarchical reinforcement learning
Carlos Florensa, Yan Duan, and Pieter Abbeel · 2017
Cited alongside, same era.
Gotta learn fast: A new benchmark for generalization in rl
Alex Nichol, Vicki Pfau, Christopher Hesse, Oleg Klimov, and John Schulman · 2018
Later among the works it cites.
Computational theories of curiosity-driven learning
Pierre-Yves Oudeyer · 2018
Later among the works it cites.
Efficient neural architecture search via parameter sharing
Hieu Pham, Melody Y Guan, Barret Zoph, Quoc V Le, and Jeff Dean · 2018
Later among the works it cites.
Some considerations on learning to explore via meta-reinforcement learning
Bradly C Stadie, Ge Yang, Rein Houthooft, Xi Chen, Yan Duan, Yuhuai Wu, Pieter Abbeel, and Ilya Sutskever · 2018
Later among the works it cites.
Evolving simple programs for playing atari games
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Meire Fortunato, Mohammad Gheshlaghi Azar, Bilal Piot, Jacob Menick, Ian Osband, Alex Graves, Vlad Mnih, Remi Munos, Demis Hassabis, Olivier Pietquin, et al · 2017
Cited alongside, same era.
Ex2: Exploration with exemplar models for deep reinforcement learning
Justin Fu, John Co-Reyes, and Sergey Levine · 2017
Cited alongside, same era.
Multi-task learning in atari video games with emergent tangled program graphs
Stephen Kelly and Malcolm I Heywood · 2017
Cited alongside, same era.
Automatic differentiation in PyTorch
Adam Paszke, Sam Gross, and Adam Lerer · 2017
Cited alongside, same era.
Curiosity-driven exploration by self-supervised prediction
Deepak Pathak, Pulkit Agrawal, Alexei A Efros, and Trevor Darrell · 2017
Cited alongside, same era.
Searching for activation functions
Prajit Ramachandran, Barret Zoph, and Quoc V Le · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
# exploration: A study of count-based exploration for deep reinforcement learning
Haoran Tang, Rein Houthooft, Davis Foote, Adam Stooke, OpenAI Xi Chen, Yan Duan, John Schulman, Filip DeTurck, and Pieter Abbeel · 2017
Cited alongside, same era.
Dennis G Wilson, Sylvain Cussat-Blanc, Hervé Luga, and Julian F Miller · 2018
Later among the works it cites.
Meta-gradient reinforcement learning
Zhongwen Xu, Hado P van Hasselt, and David Silver · 2018
Later among the works it cites.
On learning intrinsic rewards for policy gradient methods
Zeyu Zheng, Junhyuk Oh, and Satinder Singh · 2018
Later among the works it cites.
Learning transferable architectures for scalable image recognition
Barret Zoph, Vijay Vasudevan, Jonathon Shlens, and Quoc V Le · 2018
Later among the works it cites.
Neural relational inference with fast modular meta-learning
Ferran Alet, Erica Weng, Tomas Lozano-Perez, and Leslie Kaelbling · 2019
Later among the works it cites.
Mohammad Gheshlaghi Azar, Bilal Piot, Bernardo Avila Pires, Jean-Bastian Grill, Florent Altché, and Rémi Munos · 2019
Later among the works it cites.
Learning navigation behaviors end-to-end with autorl
Hao-Tien Lewis Chiang, Aleksandra Faust, Marek Fiser, and Anthony Francis · 2019
Later among the works it cites.
Learning to adapt: Meta-learning for model-based control
Ignasi Clavera, Anusha Nagabandi, Ronald S Fearing, Pieter Abbeel, Sergey Levine, and Chelsea Finn · 2019
Later among the works it cites.
Jeff Clune · 2019
Later among the works it cites.
Evolving rewards to automate reinforcement learning
Aleksandra Faust, Anthony Francis, and Dar Mehta · 2019
Later among the works it cites.
Weight agnostic neural networks
Adam Gaier and David Ha · 2019
Later among the works it cites.
Improved training speed, accuracy, and data utilization through loss function optimization
Santiago Gonzalez and Risto Miikkulainen · 2019
Later among the works it cites.
Improving generalization in meta reinforcement learning using learned objectives
Louis Kirsch, Sjoerd van Steenkiste, and Jürgen Schmidhuber · 2019
Later among the works it cites.
Self-supervised exploration via disagreement
Deepak Pathak, Dhiraj Gandhi, and Abhinav Gupta · 2019
Later among the works it cites.
Learning compositional neural programs with recursive tree search and planning
Thomas Pierrot, Guillaume Ligner, Scott Reed, Olivier Sigaud, Nicolas Perrin, Alexandre Laterre, David Kas, Karim Beguir, and Nando de Freitas · 2019
Later among the works it cites.
Few-shot bayesian imitation learning with logic over programs
Tom Silver, Kelsey R Allen, Alex K Lew, Leslie Pack Kaelbling, and Josh Tenenbaum · 2019
Later among the works it cites.
Benchmarking bonus-based exploration methods on the arcade learning environment
Adrien Ali Taïga, William Fedus, Marlos C Machado, Aaron Courville, and Marc G Bellemare · 2019
Later among the works it cites.
Discovery of useful questions as auxiliary tasks
Vivek Veeriah, Matteo Hessel, Zhongwen Xu, Richard Lewis, Janarthanan Rajendran, Junhyuk Oh, Hado van Hasselt, David Silver, and Satinder Singh · 2019
Later among the works it cites.
Rui Wang, Joel Lehman, Jeff Clune, and Kenneth O Stanley · 2019
Later among the works it cites.
Meta-world: A benchmark and evaluation for multi-task and meta-reinforcement learning, 2019
Tianhe Yu, Deirdre Quillen, Zhanpeng He, Ryan Julian, Karol Hausman, Sergey Levine, and Chelsea Finn · 2019
Later among the works it cites.
Evolving loss functions with multivariate taylor polynomial parameterizations, 2020
Santiago Gonzalez and Risto Miikkulainen · 2020
Closest in time.