Fetching the paper…
Reading the bibliography…
Self-tuning algorithms that adapt the learning process online encourage more effective and robust learning.
On the spectral theorem for normal operators
R. G. Douglas and Carl Pearcy · 1970
Earlier work this paper cites.
Neuro-dynamic programming
D. P. Bertsekas and J. N. Tsitsiklis · 1996
Earlier work this paper cites.
Actor-critic algorithms
Vijay Konda and John Tsitsiklis · 2000
Earlier work this paper cites.
Model-agnostic meta-learning for fast adaptation of deep networks, 2017
Chelsea Finn, Pieter Abbeel, and Sergey Levine · 2017
Earlier work this paper cites.
JAX: composable transformations of Python+NumPy programs, 2018
James Bradbury, Roy Frostig, Peter Hawkins, Matthew James Johnson, Chris Leary, Dougal Maclaurin, George Necula, Adam Paszke, Jake VanderPlas, Skye Wanderman-Milne, and Qiao Zhang · 2018
Earlier work this paper cites.
Meta-reinforcement learning of structured exploration strategies, 2018
Abhishek Gupta, Russell Mendonca, YuXuan Liu, Pieter Abbeel, and Sergey Levine · 2018
Earlier work this paper cites.
Evolved policy gradients, 2018
Rein Houthooft, Richard Y. Chen, Phillip Isola, Bradly C. Stadie, Filip Wolski, Jonathan Ho, and Pieter Abbeel · 2018
Earlier work this paper cites.
On first-order meta-learning algorithms, 2018
Alex Nichol, Joshua Achiam, and John Schulman · 2018
Earlier work this paper cites.
High-dimensional continuous control using generalized advantage estimation, 2018
John Schulman, Philipp Moritz, Sergey Levine, Michael Jordan, and Pieter Abbeel · 2018
Earlier work this paper cites.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Cited alongside, same era.
Understanding short-horizon bias in stochastic meta-optimization, 2018
Yuhuai Wu, Mengye Ren, Renjie Liao, and Roger Grosse · 2018
Cited alongside, same era.
Meta-gradient reinforcement learning, 2018
Zhongwen Xu, Hado van Hasselt, and David Silver · 2018
Cited alongside, same era.
On learning intrinsic rewards for policy gradient methods, 2018
Zeyu Zheng, Junhyuk Oh, and Satinder Singh · 2018
Cited alongside, same era.
Understanding and correcting pathologies in the training of learned optimizers, 2019
Luke Metz, Niru Maheswaranathan, Jeremy Nixon, C. Daniel Freeman, and Jascha Sohl-Dickstein · 2019
Cited alongside, same era.
Efficient off-policy meta-reinforcement learning via probabilistic context variables, 2019
Kate Rakelly, Aurick Zhou, Deirdre Quillen, Chelsea Finn, and Sergey Levine · 2019
Improving generalization in meta reinforcement learning using learned objectives, 2020
Louis Kirsch, Sjoerd van Steenkiste, and Jürgen Schmidhuber · 2020
Later among the works it cites.
Beyond exponentially discounted sum: Automatic learning of return function, 2020
Yufei Wang, Qiwei Ye, and Tie-Yan Liu · 2020
Later among the works it cites.
Meta-gradient reinforcement learning with an objective discovered online, 2020
Zhongwen Xu, Hado van Hasselt, Matteo Hessel, Junhyuk Oh, Satinder Singh, and David Silver · 2020
Later among the works it cites.
Online meta-critic learning for off-policy actor-critic methods, 2020
Wei Zhou, Yiying Li, Yongxin Yang, Huaimin Wang, and Timothy M. Hospedales · 2020
Later among the works it cites.
Varibad: A very good method for bayes-adaptive deep rl via meta-learning, 2020
Luisa Zintgraf, Kyriacos Shiarlis, Maximilian Igl, Sebastian Schulze, Yarin Gal, Katja Hofmann, and Shimon Whiteson · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Discovery of useful questions as auxiliary tasks, 2019
Vivek Veeriah, Matteo Hessel, Zhongwen Xu, Richard Lewis, Janarthanan Rajendran, Junhyuk Oh, Hado van Hasselt, David Silver, and Satinder Singh · 2019
Cited alongside, same era.
Fast context adaptation via meta-learning, 2019
Luisa M Zintgraf, Kyriacos Shiarlis, Vitaly Kurin, Katja Hofmann, and Shimon Whiteson · 2019
Cited alongside, same era.
Discount factor as a regularizer in reinforcement learning, 2020
Ron Amit, Ron Meir, and Kamil Ciosek · 2020
Cited alongside, same era.
Meta-learning via learned loss, 2021
Sarah Bechtle, Artem Molchanov, Yevgen Chebotar, Edward Grefenstette, Ludovic Righetti, Gaurav Sukhatme, and Franziska Meier · 2021
Closest in time.
Bootstrapped meta-learning, 2021
Sebastian Flennerhag, Yannick Schroecker, Tom Zahavy, Hado van Hasselt, David Silver, and Satinder Singh · 2021
Closest in time.
Discovering reinforcement learning algorithms, 2021
Junhyuk Oh, Matteo Hessel, Wojciech M. Czarnecki, Zhongwen Xu, Hado van Hasselt, Satinder Singh, and David Silver · 2021
Closest in time.
A self-tuning actor-critic algorithm, 2021
Tom Zahavy, Zhongwen Xu, Vivek Veeriah, Matteo Hessel, Junhyuk Oh, Hado van Hasselt, David Silver, and Satinder Singh · 2021
Closest in time.