Fetching the paper…
Reading the bibliography…
We consider Model-Agnostic Meta-Learning (MAML) methods for Reinforcement Learning (RL) problems, where the goal is to find a policy using data from several tasks represented by Markov Decision Processes (MDPs) that can be updated by one step of stochastic policy gradient for the realized MDP.
R. J. Williams, “Simple statistical gradient-following algorithms for connectionist reinforcement learning,” Machine learning
1992
Earlier work this paper cites.
Springer, 2004
Y. Nesterov, Introductory Lectures on Convex Optimization: A Basic Course · 2004
Earlier work this paper cites.
J. Peters and S. Schaal, “Reinforcement learning of motor skills with policy gradients,” Neural networks
2008
Earlier work this paper cites.
E. Todorov, T. Erez, and Y. Tassa, “Mujoco: A physics engine for model-based control,” in 2012 IEEE/RSJ International Conference on Intelligent Robots and Systems
2012
Earlier work this paper cites.
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
C. Finn, P. Abbeel, and S. Levine, “Model-agnostic meta-learning for fast adaptation of deep networks,” in Proceedings of the 34th International Conference on Machine Learning
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
J. X. Wang, Z. Kurth-Nelson, D. Kumaran, D. Tirumala, H. Soyer, J. Z. Leibo, D. Hassabis, and M. Botvinick, “Prefrontal cortex as a meta-reinforcement learning system,” Nature neuroscience
2018
Earlier work this paper cites.
2018
Cited alongside, same era.
A. Gupta, R. Mendonca, Y. Liu, P. Abbeel, and S. Levine, “Meta-reinforcement learning of structured exploration strategies,” in Advances in Neural Information Processing Systems
2018
Cited alongside, same era.
2018
Cited alongside, same era.
MIT press, 2018
R. S. Sutton and A. G. Barto, Reinforcement learning: An introduction · 2018
Cited alongside, same era.
M. Papini, D. Binaghi, G. Canonaco, M. Pirotta, and M. Restelli, “Stochastic variance-reduced policy gradient,” in Proceedings of the 35th International Conference on Machine Learning
R. Mendonca, A. Gupta, R. Kralev, P. Abbeel, S. Levine, and C. Finn, “Guided meta-policy search,” in Advances in Neural Information Processing Systems
2019
Later among the works it cites.
A. Rajeswaran, C. Finn, S. M. Kakade, and S. Levine, “Meta-learning with implicit gradients,” in Advances in Neural Information Processing Systems
2019
Later among the works it cites.
C. Finn, A. Rajeswaran, S. Kakade, and S. Levine, “Online meta-learning,” in Proceedings of the 36th International Conference on Machine Learning
2019
Later among the works it cites.
M. Khodak, M.-F. Balcan, and A. Talwalkar, “Provable guarantees for gradient-based meta-learning,” in Proceedings of the 36th International Conference on Machine Learning
2019
Later among the works it cites.
M. Khodak, M.-F. F. Balcan, and A. S. Talwalkar, “Adaptive gradient-based meta-learning methods,” in Advances in Neural Information Processing Systems
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2018
Cited alongside, same era.
A. Reuther, J. Kepner, C. Byun, S. Samsi, W. Arcand, D. Bestor, B. Bergeron, V. Gadepally, M. Houle, M. Hubbell, et al
2018
Cited alongside, same era.
B. Stadie, G. Yang, R. Houthooft, P. Chen, Y. Duan, Y. Wu, P. Abbeel, and I. Sutskever, “The importance of sampling inmeta-reinforcement learning,” in Advances in Neural Information Processing Systems
2018
Cited alongside, same era.
2019
Cited alongside, same era.
2019
Cited alongside, same era.
H. Liu, R. Socher, and C. Xiong, “Taming maml: Efficient unbiased meta-reinforcement learning,” in International Conference on Machine Learning
2019
Cited alongside, same era.
2019
Later among the works it cites.
Z. Shen, A. Ribeiro, H. Hassani, H. Qian, and C. Mi, “Hessian aided policy gradient,” in International Conference on Machine Learning
2019
Later among the works it cites.
2019
Later among the works it cites.
A. Fallah, A. Mokhtari, and A. Ozdaglar, “On the convergence theory of gradient-based model-agnostic meta-learning algorithms,” in International Conference on Artificial Intelligence and Statistics
2020
Closest in time.
2020
Closest in time.
Y. Hu, S. Zhang, X. Chen, and N. He, “Biased stochastic first-order methods for conditional stochastic optimization and applications in meta learning,” Advances in Neural Information Processing Systems
2020
Closest in time.