Fetching the paper…
Reading the bibliography…
Gradient-based Meta-RL (GMRL) refers to methods that maintain two-level optimisation procedures wherein the outer-loop meta-learner guides the inner-loop gradient-based reinforcement learner to achieve fast adaptations.
Stochastic games
Lloyd S Shapley · 1953
Earlier work this paper cites.
Evolutionary principles in self-referential learning, or on learning how to learn: the meta-meta-… hook
Jürgen Schmidhuber · 1987
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
Rl: Fast reinforcement learning via slow reinforcement learning
Yan Duan, John Schulman, Xi Chen, Peter L Bartlett, Ilya Sutskever, and Pieter Abbeel · 2016
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Earlier work this paper cites.
Learning to reinforcement learn
Jane X Wang, Zeb Kurth-Nelson, Dhruva Tirumala, Hubert Soyer, Joel Z Leibo, Remi Munos, Charles Blundell, Dharshan Kumaran, and Matt Botvinick · 2016
Earlier work this paper cites.
Continuous adaptation via meta-learning in nonstationary and competitive environments
Maruan Al-Shedivat, Trapit Bansal, Yuri Burda, Ilya Sutskever, Igor Mordatch, and Pieter Abbeel · 2017
Earlier work this paper cites.
Online learning rate adaptation with hypergradient descent
Atilim Gunes Baydin, Robert Cornish, David Martinez Rubio, Mark Schmidt, and Frank Wood · 2017
Earlier work this paper cites.
Model-agnostic meta-learning for fast adaptation of deep networks
Chelsea Finn, Pieter Abbeel, and Sergey Levine · 2017
Earlier work this paper cites.
Learning with opponent-learning awareness
Jakob N Foerster, Richard Y Chen, Maruan Al-Shedivat, Shimon Whiteson, Pieter Abbeel, and Igor Mordatch · 2017
Earlier work this paper cites.
Forward and reverse gradient-based hyperparameter optimization
Luca Franceschi, Michele Donini, Paolo Frasconi, and Massimiliano Pontil · 2017
Earlier work this paper cites.
Meta-reinforcement learning of structured exploration strategies
Abhishek Gupta, Russell Mendonca, YuXuan Liu, Pieter Abbeel, and Sergey Levine · 2018
Earlier work this paper cites.
Pytorch implementations of reinforcement learning algorithms
Ilya Kostrikov · 2018
Earlier work this paper cites.
Stable opponent shaping in differentiable games
Alistair Letcher, Jakob Foerster, David Balduzzi, Tim Rocktäschel, and Shimon Whiteson · 2018
Earlier work this paper cites.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Cited alongside, same era.
Meta-gradient reinforcement learning
Zhongwen Xu, Hado van Hasselt, and David Silver · 2018
Cited alongside, same era.
On learning intrinsic rewards for policy gradient methods
Zeyu Zheng, Junhyuk Oh, and Satinder Singh · 2018
Cited alongside, same era.
Gregory Farquhar, Shimon Whiteson, and Jakob Foerster · 2019
Cited alongside, same era.
Improving generalization in meta reinforcement learning using learned objectives
Louis Kirsch, Sjoerd van Steenkiste, and Juergen Schmidhuber · 2019
Cited alongside, same era.
A policy gradient algorithm for learning to learn in multiagent reinforcement learning
Dong-Ki Kim, Miao Liu, Matthew Riemer, Chuangchuang Sun, Marwa Abdulhai, Golnaz Habibi, Sebastian Lopez-Cot, Gerald Tesauro, and Jonathan P How · 2020
Later among the works it cites.
Discovering reinforcement learning algorithms
Junhyuk Oh, Matteo Hessel, Wojciech M Czarnecki, Zhongwen Xu, Hado van Hasselt, Satinder Singh, and David Silver · 2020
Later among the works it cites.
Meta-gradient reinforcement learning with an objective discovered online
Zhongwen Xu, Hado van Hasselt, Matteo Hessel, Junhyuk Oh, Satinder Singh, and David Silver · 2020
Later among the works it cites.
A self-tuning actor-critic algorithm
Tom Zahavy, Zhongwen Xu, Vivek Veeriah, Matteo Hessel, Junhyuk Oh, Hado van Hasselt, David Silver, and Satinder Singh · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Taming maml: Efficient unbiased meta-reinforcement learning
Hao Liu, Richard Socher, and Caiming Xiong · 2019
Cited alongside, same era.
A baseline for any order gradient estimation in stochastic computation graphs
Jingkai Mao, Jakob Foerster, Tim Rocktäschel, Maruan Al-Shedivat, Gregory Farquhar, and Shimon Whiteson · 2019
Cited alongside, same era.
Efficient off-policy meta-reinforcement learning via probabilistic context variables
Kate Rakelly, Aurick Zhou, Chelsea Finn, Sergey Levine, and Deirdre Quillen · 2019
Cited alongside, same era.
Promp: Proximal meta-policy search
Jonas Rothfuss, Dennis Lee, Ignasi Clavera, Tamim Asfour, and Pieter Abbeel · 2019
Cited alongside, same era.
Discovery of useful questions as auxiliary tasks
Vivek Veeriah, Matteo Hessel, Zhongwen Xu, Richard Lewis, Janarthanan Rajendran, Junhyuk Oh, Hado van Hasselt, David Silver, and Satinder Singh · 2019
Cited alongside, same era.
On the convergence theory of gradient-based model-agnostic meta-learning algorithms
Alireza Fallah, Aryan Mokhtari, and Asuman Ozdaglar · 2020
Cited alongside, same era.
Towards effective context for meta-reinforcement learning: an approach based on contrastive learning
Haotian Fu, Hongyao Tang, Jianye Hao, Chen Chen, Xidong Feng, Dong Li, and Wulong Liu · 2020
Cited alongside, same era.
What can learned intrinsic rewards capture?
Zeyu Zheng, Junhyuk Oh, Matteo Hessel, Zhongwen Xu, Manuel Kroiss, Hado Van Hasselt, David Silver, and Satinder Singh · 2020
Later among the works it cites.
Online meta-critic learning for off-policy actor-critic methods
Wei Zhou, Yiying Li, Yongxin Yang, Huaimin Wang, and Timothy Hospedales · 2020
Later among the works it cites.
Meta learning via learned loss
Sarah Bechtle, Artem Molchanov, Yevgen Chebotar, Edward Grefenstette, Ludovic Righetti, Gaurav Sukhatme, and Franziska Meier · 2021
Closest in time.
One step at a time: Pros and cons of multi-step meta-gradient reinforcement learning
Clément Bonnet, Paul Caron, Thomas Barrett, Ian Davies, and Alexandre Laterre · 2021
Closest in time.
On the convergence theory of debiased model-agnostic meta-reinforcement learning
Alireza Fallah, Kristian Georgiev, Aryan Mokhtari, and Asuman Ozdaglar · 2021
Closest in time.
Neural auto-curricula in two-player zero-sum games
Xidong Feng, Oliver Slumbers, Ziyu Wan, Bo Liu, Stephen McAleer, Ying Wen, Jun Wang, and Yaodong Yang · 2021
Closest in time.
Unifying gradient estimators for meta-reinforcement learning via off-policy evaluation
Yunhao Tang, Tadashi Kozuno, Mark Rowland, Remi Munos, and Michal Valko · 2021
Closest in time.
Discovery of options via meta-learned subgoals
Vivek Veeriah, Tom Zahavy, Matteo Hessel, Zhongwen Xu, Junhyuk Oh, Iurii Kemaev, Hado van Hasselt, David Silver, and Satinder Singh · 2021
Closest in time.
No DICE: An investigation of the bias-variance tradeoff in meta-gradients
Risto Vuorio, Jacob Austin Beck, Gregory Farquhar, Jakob Nicolaus Foerster, and Shimon Whiteson · 2021
Closest in time.
Torchopt: An efficient library for differentiable optimization
Jie Ren, Xidong Feng, Bo Liu, Xuehai Pan, Yao Fu, Luo Mai, and Yaodong Yang · 2022
Closest in time.