Fetching the paper…
Reading the bibliography…
Meta-Reinforcement learning approaches aim to develop learning procedures that can adapt quickly to a distribution of tasks with the help of a few examples.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Earlier work this paper cites.
Rl 2 : Fast reinforcement learning via slow reinforcement learning
Y. Duan, J. Schulman, X. Chen, P. L. Bartlett, I. Sutskever, and P. Abbeel · 2016
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu · 2016
Earlier work this paper cites.
Benchmarking deep reinforcement learning for continuous control
Y. Duan, X. Chen, R. Houthooft, J. Schulman, and P. Abbeel · 2016
Earlier work this paper cites.
Deep successor reinforcement learning
T. D. Kulkarni, A. Saeedi, S. Gautam, and S. J. Gershman · 2016
Earlier work this paper cites.
Mastering the game of go without human knowledge
D. Silver, J. Schrittwieser, K. Simonyan, I. Antonoglou, A. Huang, A. Guez, T. Hubert, L. Baker, M. Lai, A. Bolton, et al · 2017
Earlier work this paper cites.
Molecular de-novo design through deep reinforcement learning
M. Olivecrona, T. Blaschke, O. Engkvist, and H. Chen · 2017
Earlier work this paper cites.
Model-agnostic meta-learning for fast adaptation of deep networks
C. Finn, P. Abbeel, and S. Levine · 2017
Earlier work this paper cites.
C. Finn and S. Levine · 2017
Earlier work this paper cites.
Prototypical networks for few-shot learning
J. Snell, K. Swersky, and R. Zemel · 2017
Cited alongside, same era.
Self-supervision for reinforcement learning
P. Mahmoudieh · 2017
Cited alongside, same era.
Curiosity-driven exploration by self-supervised prediction
D. Pathak, P. Agrawal, A. A. Efros, and T. Darrell · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Cited alongside, same era.
A study on overfitting in deep reinforcement learning
C. Zhang, O. Vinyals, R. Munos, and S. Bengio · 2018
Cited alongside, same era.
Some considerations on learning to explore via meta-reinforcement learning
Meta-reinforcement learning of structured exploration strategies
A. Gupta, R. Mendonca, Y. Liu, P. Abbeel, and S. Levine · 2018
Later among the works it cites.
Self-supervised deep reinforcement learning with generalized computation graphs for robot navigation
G. Kahn, A. Villaflor, B. Ding, P. Abbeel, and S. Levine · 2018
Later among the works it cites.
Learning to learn how to learn: Self-adaptive visual navigation using meta-learning
M. Wortsman, K. Ehsani, M. Rastegari, A. Farhadi, and R. Mottaghi · 2018
Later among the works it cites.
Dice: The infinitely differentiable monte-carlo estimator
J. Foerster, G. Farquhar, M. Al-Shedivat, T. Rocktäschel, E. P. Xing, and S. Whiteson · 2018
Later among the works it cites.
Caml: Fast context adaptation via meta-learning
L. M. Zintgraf, K. Shiarlis, V. Kurin, K. Hofmann, and S. Whiteson · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
B. C. Stadie, G. Yang, R. Houthooft, X. Chen, Y. Duan, Y. Wu, P. Abbeel, and I. Sutskever · 2018
Cited alongside, same era.
Promp: Proximal meta-policy search
J. Rothfuss, D. Lee, I. Clavera, T. Asfour, and P. Abbeel · 2018
Cited alongside, same era.
Learning to compare: Relation network for few-shot learning
F. Sung, Y. Yang, L. Zhang, T. Xiang, P. H. Torr, and T. M. Hospedales · 2018
Cited alongside, same era.
On first-order meta-learning algorithms
A. Nichol, J. Achiam, and J. Schulman · 2018
Cited alongside, same era.
Later among the works it cites.
Meta-learning for contextual bandit exploration
A. Sharaf and H. Daumé III · 2019
Closest in time.
Efficient off-policy meta-reinforcement learning via probabilistic context variables
K. Rakelly, A. Zhou, D. Quillen, C. Finn, and S. Levine · 2019
Closest in time.
Self-supervised learning of image embedding for continuous control
C. Florensa, J. Degrave, N. Heess, J. T. Springenberg, and M. Riedmiller · 2019
Closest in time.
Norml: No-reward meta learning
Y. Yang, K. Caluwaerts, A. Iscen, J. Tan, and C. Finn · 2019
Closest in time.