Fetching the paper…
Reading the bibliography…
We propose a method for meta-learning reinforcement learning algorithms by searching over the space of computational graphs which compute the loss function for a value-based model-free RL agent to optimize.
Evolutionary principles in self-referential learning
Juergen Schmidhuber · 1987
Earlier work this paper cites.
Learning a synaptic learning rule
Yoshua Bengio, S. Bengio, and J. Cloutier · 1991
Earlier work this paper cites.
A comparative analysis of selection schemes used in genetic algorithms
David E Goldberg and Kalyanmoy Deb · 1991
Earlier work this paper cites.
Adaptation in Natural and Artificial Systems
John H. Holland · 1992
Earlier work this paper cites.
Genetic programming - on the programming of computers by means of natural selection
John Koza · 1993
Earlier work this paper cites.
A self-referential weight matrix
Juergen Schmidhuber · 1993
Earlier work this paper cites.
Use of genetic programming for the search of a new learning rule for neural networks
S. Bengio, Yoshua Bengio, and J. Cloutier · 1994
Earlier work this paper cites.
Evolving neural networks through augmenting topologies
Kenneth O Stanley and Risto Miikkulainen · 2002
Earlier work this paper cites.
Meta-learning curiosity algorithms
Ferran Alet, M. F. Schneider, Tomas Lozano-Perez, and L. Kaelbling · 2003
Earlier work this paper cites.
Synthesis of interest point detectors through genetic programming
L. Trujillo and G. Olague · 2006
Earlier work this paper cites.
Evolutionary function approximation for reinforcement learning
S. Whiteson and P. Stone · 2006
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin A. Riedmiller · 2013
Earlier work this paper cites.
Neural programmer-interpreters
Scott Reed and Nando De Freitas · 2015
Earlier work this paper cites.
Deep reinforcement learning with double q-learning
Hado van Hasselt, Arthur Guez, and David Silver · 2015
Earlier work this paper cites.
Rl2: Fast reinforcement learning via slow reinforcement learning
Yan Duan, John Schulman, Xi Chen, Peter L. Bartlett, Ilya Sutskever, and P. Abbeel · 2016
Earlier work this paper cites.
Neural architecture search with reinforcement learning
Barret Zoph and Quoc V Le · 2016
Earlier work this paper cites.
Neural optimizer search with reinforcement learning
Irwan Bello, Barret Zoph, Vijay Vasudevan, and Quoc V Le · 2017
Cited alongside, same era.
Model-agnostic meta-learning for fast adaptation of deep networks
Chelsea Finn, P. Abbeel, and S. Levine · 2017
Cited alongside, same era.
Progressive neural architecture search
Chenxi Liu, Barret Zoph, Jonathon Shlens, Wei Hua, Li-Jia Li, Li Fei-Fei, Alan L. Yuille, Jonathan Huang, and Kevin Murphy · 2017
Cited alongside, same era.
Large-scale evolution of image classifiers
Esteban Real, Sherry Moore, Andrew Selle, Saurabh Saxena, Yutaka Leon Suematsu, Jie Tan, Quoc V. Le, and Alexey Kurakin · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Efficient neural architecture search via parameter sharing
Hieu Pham, Melody Y Guan, Barret Zoph, Quoc V Le, and Jeff Dean · 2018
Later among the works it cites.
Learning transferable architectures for scalable image recognition
Barret Zoph, Vijay Vasudevan, Jonathon Shlens, and Quoc V Le · 2018
Later among the works it cites.
Meta-learning via learned loss
Yevgen Chebotar, Artem Molchanov, Sarah Bechtle, Ludovic Righetti, F. Meier, and Gaurav S. Sukhatme · 2019
Later among the works it cites.
Evolving rewards to automate reinforcement learning
Aleksandra Faust, Anthony Francis, and Dar Mehta · 2019
Later among the works it cites.
Learning compositional neural programs with recursive tree search and planning
Thomas Pierrot, Guillaume Ligner, Scott Reed, Olivier Sigaud, Nicolas Perrin, Alexandre Laterre, David Kas, Karim Beguir, and Nando de Freitas · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Learning to reinforcement learn
Jane X. Wang, Zeb Kurth-Nelson, Hubert Soyer, Joel Z. Leibo, Dhruva Tirumala, Rémi Munos, Charles Blundell, D. Kumaran, and Matt M. Botvinick · 2017
Cited alongside, same era.
Efficient architecture search by network transformation
Han Cai, Tianyao Chen, Weinan Zhang, Yong Yu, and Jun Wang · 2018
Cited alongside, same era.
Minimalistic gridworld environment for openai gym
Maxime Chevalier-Boisvert, Lucas Willems, and Suman Pal · 2018
Cited alongside, same era.
Neural architecture search: A survey
Thomas Elsken, Jan Hendrik Metzen, and Frank Hutter · 2018
Cited alongside, same era.
Chelsea Finn and S. Levine · 2018
Cited alongside, same era.
Addressing function approximation error in actor-critic methods
Scott Fujimoto, H. V. Hoof, and David Meger · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Toumas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Cited alongside, same era.
Regularized evolution for image classifier architecture search
Esteban Real, Alok Aggarwal, Yanping Huang, and Quoc V. Le · 2019
Later among the works it cites.
D. So, Chen Liang, and Quoc V. Le · 2019
Later among the works it cites.
Sample-efficient automated deep reinforcement learning
Jörg K. H. Franke, Gregor Köhler, André Biedenkapp, and Frank Hutter · 2020
Later among the works it cites.
Improving generalization in meta reinforcement learning using learned objectives
Louis Kirsch, Sjoerd van Steenkiste, and J. Schmidhuber · 2020
Later among the works it cites.
Conservative q-learning for offline reinforcement learning
Aviral Kumar, Aurick Zhou, G. Tucker, and Sergey Levine · 2020
Later among the works it cites.
Discovering reinforcement learning algorithms
Junhyuk Oh, Matteo Hessel, Wojciech Marian Czarnecki, Zhongwen Xu, Hado van Hasselt, Satinder Singh, and David Silver · 2020
Later among the works it cites.
Automl-zero: Evolving machine learning algorithms from scratch
Esteban Real, Chen Liang, David So, and Quoc V. Le · 2020
Later among the works it cites.
Reinforcement learning with chromatic networks for compact architecture search
Xingyou Song, Krzysztof Choromanski, Jack Parker-Holder, Yunhao Tang, Wenbo Gao, Aldo Pacchiano, Tamas Sarlos, Deepali Jain, and Yuxiang Yang · 2020
Later among the works it cites.
Online hyper-parameter tuning in off-policy learning via evolutionary strategies, 2020
Yunhao Tang and Krzysztof Choromanski · 2020
Later among the works it cites.
Munchausen reinforcement learning
Nino Vieillard, Olivier Pietquin, and M. Geist · 2020
Later among the works it cites.