Fetching the paper…
Reading the bibliography…
Gradient-based meta-learners such as Model-Agnostic Meta-Learning (MAML) have shown strong few-shot performance in supervised and reinforcement learning settings.
Efficient Off-Policy Meta-Reinforcement Learning via Probabilistic Context Variables
K. Rakelly, A. Zhou, D. Quillen, C. Finn, and S. Levine · 1903
Earlier work this paper cites.
B. Mehta, M. Diaz, F. Golemo, C. J. Pal, and L. Paull · 1904
Earlier work this paper cites.
Learning domain randomization distributions for transfer of locomotion policies
M. Mozifian, J. C. G. Higuera, D. Meger, and G. Dudek · 1906
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
R. J. Williams · 1992
Earlier work this paper cites.
Curriculum learning
Y. Bengio, J. Louradour, R. Collobert, and J. Weston · 2009
Earlier work this paper cites.
Modeling Purposeful Adaptive Behavior with the Principle of Maximum Causal Entropy
B. D. Ziebart · 2010
Earlier work this paper cites.
Siamese Neural Networks for One-shot Image Recognition
G. Koch, R. Zemel, and R. Salakhutdinov · 2015
Earlier work this paper cites.
Trust Region Policy Optimization, 2015
J. Schulman, S. Levine, P. Moritz, M. I. Jordan, and P. Abbeel · 2015
Earlier work this paper cites.
RL2: Fast Reinforcement Learning via Slow Reinforcement Learning
Y. Duan, J. Schulman, X. Chen, P. L. Bartlett, I. Sutskever, and P. Abbeel · 2016
Earlier work this paper cites.
Meta-Learning with Memory-Augmented Neural Networks
A. Santoro, S. Bartunov, M. Botvinick, D. Wierstra, and T. Lillicrap · 2016
Earlier work this paper cites.
Learning the curriculum with Bayesian optimization for task-specific word representation learning
Y. Tsvetkov, M. Faruqui, W. Ling, B. MacWhinney, and C. Dyer · 2016
Earlier work this paper cites.
Matching Networks for One Shot Learning
O. Vinyals, C. Blundell, T. P. Lillicrap, K. Kavukcuoglu, and D. Wierstra · 2016
Earlier work this paper cites.
Learning to reinforcement learn
J. X. Wang, Z. Kurth-Nelson, D. Tirumala, H. Soyer, J. Z. Leibo, R. Munos, C. Blundell, D. Kumaran, and M. Botvinick · 2016
Cited alongside, same era.
Reverse curriculum generation for reinforcement learning, 2017
C. Florensa, D. Held, M. Wulfmeier, M. Zhang, and P. Abbeel · 2017
Cited alongside, same era.
Automated curriculum learning for neural networks, 2017
A. Graves, M. G. Bellemare, J. Menick, R. Munos, and K. Kavukcuoglu · 2017
Cited alongside, same era.
Emergence of locomotion behaviours in rich environments, 2017
N. Heess, D. TB, S. Sriram, J. Lemmon, J. Merel, G. Wayne, Y. Tassa, T. Erez, Z. Wang, S. M. A. Eslami, M. Riedmiller, and D. Silver · 2017
Cited alongside, same era.
Stein variational policy gradient, 2017
Y. Liu, P. Ramachandran, Q. Liu, and J. Peng · 2017
Cited alongside, same era.
Meta-reinforcement learning of structured exploration strategies, 2018
A. Gupta, R. Mendonca, Y. Liu, P. Abbeel, and S. Levine · 2018
Later among the works it cites.
A Simple Neural Attentive Meta-Learner
N. Mishra, M. Rohaninejad, X. Chen, and P. Abbeel · 2018
Later among the works it cites.
On first-order meta-learning algorithms, 2018
A. Nichol, J. Achiam, and J. Schulman · 2018
Later among the works it cites.
Promp: Proximal meta-policy search, 2018
J. Rothfuss, D. Lee, I. Clavera, T. Asfour, and P. Abbeel · 2018
Later among the works it cites.
Some considerations on learning to explore via meta-reinforcement learning, 2018
B. C. Stadie, G. Yang, R. Houthooft, X. Chen, Y. Duan, Y. Wu, P. Abbeel, and I. Sutskever · 2018
Later among the works it cites.
Reinforcement Learning: An introduction
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
T. Munkhdalai and H. Yu · 2017
Cited alongside, same era.
Robust adversarial reinforcement learning, 2017
L. Pinto, J. Davidson, R. Sukthankar, and A. Gupta · 2017
Cited alongside, same era.
Optimization as a model for few-shot learning
S. Ravi and H. Larochelle · 2017
Cited alongside, same era.
Prototypical Networks for Few-shot Learning
J. Snell, K. Swersky, and R. S. Zemel · 2017
Cited alongside, same era.
Domain randomization for transferring deep neural networks from simulation to the real world
J. Tobin, R. Fong, A. Ray, J. Schneider, W. Zaremba, and P. Abbeel · 2017
Cited alongside, same era.
The effects of negative adaptation in Model-Agnostic Meta-Learning
T. Deleu and Y. Bengio · 2018
Cited alongside, same era.
Reinforcement Learning with Model-Agnostic Meta-Learning in PyTorch, 2018
T. Deleu and H. S. Guiroy, Simon · 2018
Cited alongside, same era.
R. S. Sutton and A. G. Barto · 2018
Later among the works it cites.
Mame : Model-agnostic meta-exploration, 2019
S. Gurumurthy, S. Kumar, and K. Sycara · 2019
Later among the works it cites.
Solving rubik’s cube with a robot hand
OpenAI, I. Akkaya, M. Andrychowicz, M. Chociej, M. Litwin, B. McGrew, A. Petron, A. Paino, M. Plappert, G. Powell, R. Ribas, J. Schneider, N. Tezak, J. Tworek, P. Welinder, L. Weng, Q.-M. Yuan, W. Zaremba, and L. Zhang · 2019
Later among the works it cites.
Skew-fit: State-covering self-supervised reinforcement learning, 2019
V. H. Pong, M. Dalal, S. Lin, A. Nair, S. Bahl, and S. Levine · 2019
Later among the works it cites.
Paired open-ended trailblazer (poet): Endlessly generating increasingly complex and diverse learning environments and their solutions, 2019
R. Wang, J. Lehman, J. Clune, and K. O. Stanley · 2019
Later among the works it cites.
Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning, 2019
T. Yu, D. Quillen, Z. He, R. Julian, K. Hausman, C. Finn, and S. Levine · 2019
Later among the works it cites.