Fetching the paper…
Reading the bibliography…
Gradient-based methods for optimisation of objectives in stochastic settings with unknown or intractable dynamics require estimators of derivatives.
The optimal reward baseline for gradient-based reinforcement learning
Lex Weaver and Nigel Tao · 2001
Earlier work this paper cites.
Gradient estimation
Michael C Fu · 2006
Earlier work this paper cites.
Bias in natural actor-critic algorithms
Philip Thomas · 2014
Earlier work this paper cites.
Approximate newton methods for policy search in markov decision processes
Thomas Furmston, Guy Lever, and David Barber · 2016
Earlier work this paper cites.
Model-agnostic meta-learning for fast adaptation of deep networks
Chelsea Finn, Pieter Abbeel, and Sergey Levine · 2017
Earlier work this paper cites.
An introduction to deep reinforcement learning
Vincent François-Lavet, Peter Henderson, Riashat Islam, Marc G Bellemare, Joelle Pineau, et al · 2018
Cited alongside, same era.
Promp: Proximal meta-policy search
Jonas Rothfuss, Dennis Lee, Ignasi Clavera, Tamim Asfour, and Pieter Abbeel · 2018
Cited alongside, same era.
Some considerations on learning to explore via meta-reinforcement learning
Bradly C Stadie, Ge Yang, Rein Houthooft, Xi Chen, Yan Duan, Yuhuai Wu, Pieter Abbeel, and Ilya Sutskever · 2018
Cited alongside, same era.
Taming maml: Efficient unbiased meta-reinforcement learning
Hao Liu, Richard Socher, and Caiming Xiong · 2019
Cited alongside, same era.
Learning with opponent-learning awareness
Jakob Foerster, Richard Y Chen, Maruan Al-Shedivat, Shimon Whiteson, Pieter Abbeel, and Igor Mordatch
Cited in the paper.
Dice: The infinitely differentiable monte carlo estimator
Jakob Foerster, Gregory Farquhar, Maruan Al-Shedivat, Tim Rocktäschel, Eric Xing, and Shimon Whiteson
Cited in the paper.
Counterfactual multi-agent policy gradients
Jakob N Foerster, Gregory Farquhar, Triantafyllos Afouras, Nantas Nardelli, and Shimon Whiteson
Cited in the paper.
Gradient estimation using stochastic computation graphs
John Schulman, Nicolas Heess, Theophane Weber, and Pieter Abbeel
Cited in the paper.
High-dimensional continuous control using generalized advantage estimation
John Schulman, Philipp Moritz, Sergey Levine, Michael Jordan, and Pieter Abbeel
Cited in the paper.
A baseline for any order gradient estimation in stochastic computation graphs
Jingkai Mao, Jakob Foerster, Tim Rocktäschel, Maruan Al-Shedivat, Gregory Farquhar, and Shimon Whiteson · 2019
Closest in time.
Monte carlo gradient estimation in machine learning
Shakir Mohamed, Mihaela Rosca, Michael Figurnov, and Andriy Mnih · 2019
Closest in time.
Credit assignment techniques in stochastic computation graphs
Théophane Weber, Nicolas Heess, Lars Buesing, and David Silver · 2019
Closest in time.
Caml: Fast context adaptation via meta-learning
Luisa M Zintgraf, Kyriacos Shiarlis, Vitaly Kurin, Katja Hofmann, and Shimon Whiteson · 2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…