Fetching the paper…
Reading the bibliography…
Context, the embedding of previous collected trajectories, is a powerful construct for Meta-Reinforcement Learning (Meta-RL) algorithms.
Data-Efficient Image Recognition with Contrastive Predictive Coding
Hénaff, O. J.; Srinivas, A.; Fauw, J. D.; Razavi, A.; Doersch, C.; Eslami, S. M. A.; and van den Oord, A. 2019 · 1905
Earlier work this paper cites.
MGHRL: Meta Goal-generation for Hierarchical Reinforcement Learning
Fu, H.; Tang, H.; Hao, J.; Liu, W.; and Chen, C. 2019 · 1909
Earlier work this paper cites.
Evolutionary principles in self-referential learning
Schmidhuber, J. 1987 · 1987
Earlier work this paper cites.
Self-Organization in a Perceptual Network
Linsker, R. 1988 · 1988
Earlier work this paper cites.
Meta-neural networks that learn by learning
Naik, D. K.; and Mammone, R. 1992 · 1992
Earlier work this paper cites.
Learning to Learn
Thrun, S.; and Pratt, L. Y. 1998 · 1998
Earlier work this paper cites.
A Simple Framework for Contrastive Learning of Visual Representations
Chen, T.; Kornblith, S.; Norouzi, M.; and Hinton, G. E. 2020 · 2002
Earlier work this paper cites.
CURL: Contrastive Unsupervised Representations for Reinforcement Learning
Srinivas, A.; Laskin, M.; and Abbeel, P. 2020 · 2004
Earlier work this paper cites.
Context-aware Dynamics Model for Generalization in Model-Based Reinforcement Learning
Lee, K.; Seo, Y.; Lee, S.; Lee, H.; and Shin, J. 2020 · 2005
Earlier work this paper cites.
Learn to Effectively Explore in Context-Based Meta-RL
Zhang, J.; Wang, J.; Hu, H.; Chen, Y.; Fan, C.; and Zhang, C. 2020 · 2006
Earlier work this paper cites.
Explore then Execute: Adapting without Rewards via Factorized Meta-Reinforcement Learning
Liu, E. Z.; Raghunathan, A.; Liang, P.; and Finn, C. 2020 · 2008
Earlier work this paper cites.
Visualizing Data using t-SNE
Maaten, L. V. D.; and Hinton, G. E. 2008 · 2008
Earlier work this paper cites.
MuJoCo: A physics engine for model-based control
Todorov, E.; Erez, T.; and Tassa, Y. 2012 · 2012
Cited alongside, same era.
VariBAD: A Very Good Method for Bayes-Adaptive Deep RL via Meta-Learning
Zintgraf, L. M.; Shiarlis, K.; Igl, M.; Schulze, S.; Gal, Y.; Hofmann, K.; and Whiteson, S. 2020 · 2012
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, V.; Kavukcuoglu, K.; Silver, D.; Rusu, A. A.; Veness, J.; Bellemare, M. G.; Graves, A.; Riedmiller, M. A.; Fidjeland, A.; Ostrovski, G.; Petersen, S.; Beattie, C.; Sadik, A.; Antonoglou, I.; King, H.; Kumaran, D.; Wierstra, D.; Legg, S.; and Hassabis, D. 2015 · 2015
Cited alongside, same era.
Trust Region Policy Optimization
Schulman, J.; Levine, S.; Abbeel, P.; Jordan, M. I.; and Moritz, P. 2015 · 2015
Cited alongside, same era.
Brockman, G.; Cheung, V.; Pettersson, L.; Schneider, J.; Schulman, J.; Tang, J.; and Zaremba, W. 2016 · 2016
Cited alongside, same era.
Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor
Haarnoja, T.; Zhou, A.; Abbeel, P.; and Levine, S. 2018 · 2018
Later among the works it cites.
Time-Contrastive Networks: Self-Supervised Learning from Video
Sermanet, P.; Lynch, C.; Chebotar, Y.; Hsu, J.; Jang, E.; Schaal, S.; and Levine, S. 2018 · 2018
Later among the works it cites.
Representation Learning with Contrastive Predictive Coding
van den Oord, A.; Li, Y.; and Vinyals, O. 2018 · 2018
Later among the works it cites.
Unsupervised Feature Learning via Non-parametric Instance Discrimination
Wu, Z.; Xiong, Y.; Yu, S.; and Lin, D. 2018 · 2018
Later among the works it cites.
Unsupervised State Representation Learning in Atari
Anand, A.; Racah, E.; Ozair, S.; Bengio, Y.; Côté, M.; and Hjelm, R. D. 2019 · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
RL$ˆ2$: Fast Reinforcement Learning via Slow Reinforcement Learning
Duan, Y.; Schulman, J.; Chen, X.; Bartlett, P. L.; Sutskever, I.; and Abbeel, P. 2016 · 2016
Cited alongside, same era.
Continuous control with deep reinforcement learning
Lillicrap, T. P.; Hunt, J. J.; Pritzel, A.; Heess, N.; Erez, T.; Tassa, Y.; Silver, D.; and Wierstra, D. 2016 · 2016
Cited alongside, same era.
Learning to reinforcement learn
Wang, J. X.; Kurth-Nelson, Z.; Tirumala, D.; Soyer, H.; Leibo, J. Z.; Munos, R.; Blundell, C.; Kumaran, D.; and Botvinick, M. 2016 · 2016
Cited alongside, same era.
Model-Agnostic Meta-Learning for Fast Adaptation of Deep Networks
Finn, C.; Abbeel, P.; and Levine, S. 2017 · 2017
Cited alongside, same era.
Curiosity-driven Exploration by Self-supervised Prediction
Pathak, D.; Agrawal, P.; Efros, A. A.; and Darrell, T. 2017 · 2017
Cited alongside, same era.
Learning Actionable Representations from Visual Observations
Dwibedi, D.; Tompson, J.; Lynch, C.; and Sermanet, P. 2018 · 2018
Cited alongside, same era.
Meta-Reinforcement Learning of Structured Exploration Strategies
Gupta, A.; Mendonca, R.; Liu, Y.; Abbeel, P.; and Levine, S. 2018 · 2018
Cited alongside, same era.
Taming MAML: Efficient unbiased meta-reinforcement learning
Liu, H.; Socher, R.; and Xiong, C. 2019 · 2019
Later among the works it cites.
On Variational Bounds of Mutual Information
Poole, B.; Ozair, S.; van den Oord, A.; Alemi, A.; and Tucker, G. 2019 · 2019
Later among the works it cites.
Efficient Off-Policy Meta-Reinforcement Learning via Probabilistic Context Variables
Rakelly, K.; Zhou, A.; Finn, C.; Levine, S.; and Quillen, D. 2019 · 2019
Later among the works it cites.
ProMP: Proximal Meta-Policy Search
Rothfuss, J.; Lee, D.; Clavera, I.; Asfour, T.; and Abbeel, P. 2019 · 2019
Later among the works it cites.
Environment Probing Interaction Policies
Zhou, W.; Pinto, L.; and Gupta, A. 2019 · 2019
Later among the works it cites.
Meta-Q-Learning
Fakoor, R.; Chaudhari, P.; Soatto, S.; and Smola, A. J. 2020 · 2020
Closest in time.
Momentum Contrast for Unsupervised Visual Representation Learning
He, K.; Fan, H.; Wu, Y.; Xie, S.; and Girshick, R. B. 2020 · 2020
Closest in time.