Fetching the paper…
Reading the bibliography…
Meta-reinforcement learning (meta-RL) aims to learn from multiple training tasks the ability to adapt efficiently to unseen test tasks.
Shift of bias for inductive concept learning
P. E. Utgoff · 1986
Earlier work this paper cites.
Evolutionary principles in self-referential learning, or on learning how to learn: the meta-meta-… hook
J. Schmidhuber · 1987
Earlier work this paper cites.
Training connectionist networks with queries and selective sampling
L. E. Atlas, D. A. Cohn, and R. E. Ladner · 1990
Earlier work this paper cites.
Integrated architectures for learning, planning, and reacting based on approximating dynamic programming
R. S. Sutton · 1990
Earlier work this paper cites.
On the optimization of a synaptic learning rule
S. Bengio, Y. Bengio, J. Cloutier, and J. Gecsei · 1992
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
R. J. Williams · 1992
Earlier work this paper cites.
A sequential algorithm for training text classifiers
D. D. Lewis and W. A. Gale · 1994
Earlier work this paper cites.
Fast exact multiplication by the hessian
B. A. Pearlmutter · 1994
Earlier work this paper cites.
Active Learning: 101 Strategies To Teach Any Subject
M. Silberman · 1996
Earlier work this paper cites.
Is learning the n-th thing any easier than learning the first?
S. Thrun · 1996
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
R. S. Sutton, D. A. McAllester, S. P. Singh, and Y. Mansour · 2000
Earlier work this paper cites.
Learning to learn using gradient descent
S. Hochreiter, A. S. Younger, and P. R. Conwell · 2001
Earlier work this paper cites.
Active learning literature survey
B. Settles · 2009
Earlier work this paper cites.
Learning to learn
S. Thrun and L. Pratt · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
E. Todorov, T. Erez, and Y. Tassa · 2012
Earlier work this paper cites.
Playing atari with deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. Graves, I. Antonoglou, D. Wierstra, and M. Riedmiller · 2013
Earlier work this paper cites.
G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, and W. Zaremba · 2016
Earlier work this paper cites.
Rl2: Fast reinforcement learning via slow reinforcement learning
Y. Duan, J. Schulman, X. Chen, P. L. Bartlett, I. Sutskever, and P. Abbeel · 2016
Earlier work this paper cites.
End-to-end training of deep visuomotor policies
S. Levine, C. Finn, T. Darrell, and P. Abbeel · 2016
Cited alongside, same era.
Epopt: Learning robust neural network policies using model ensembles
A. Rajeswaran, S. Ghotra, B. Ravindran, and S. Levine · 2016
Cited alongside, same era.
Mastering the game of go with deep neural networks and tree search
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. Van Den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, et al · 2016
Cited alongside, same era.
Learning to reinforcement learn
J. X. Wang, Z. Kurth-Nelson, D. Tirumala, H. Soyer, J. Z. Leibo, R. Munos, C. Blundell, D. Kumaran, and M. Botvinick · 2016
Cited alongside, same era.
Model-agnostic meta-learning for fast adaptation of deep networks
C. Finn, P. Abbeel, and S. Levine · 2017
Cited alongside, same era.
Bootstrapping the expressivity with model-based planning
K. Dong, Y. Luo, and T. Ma · 2019
Later among the works it cites.
R. Fakoor, P. Chaudhari, S. Soatto, and A. J. Smola · 2019
Later among the works it cites.
Meta reinforcement learning as task inference
J. Humplik, A. Galashov, L. Hasenclever, P. A. Ortega, Y. W. Teh, and N. Heess · 2019
Later among the works it cites.
When to trust your model: Model-based policy optimization
M. Janner, J. Fu, M. Zhang, and S. Levine · 2019
Later among the works it cites.
Improving generalization in meta reinforcement learning using learned objectives
L. Kirsch, S. van Steenkiste, and J. Schmidhuber · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Towards deep learning models resistant to adversarial attacks
A. Madry, A. Makelov, L. Schmidt, D. Tsipras, and A. Vladu · 2017
Cited alongside, same era.
Curiosity-driven exploration by self-supervised prediction
D. Pathak, P. Agrawal, A. A. Efros, and T. Darrell · 2017
Cited alongside, same era.
Prototypical networks for few-shot learning
J. Snell, K. Swersky, and R. Zemel · 2017
Cited alongside, same era.
Sample-efficient reinforcement learning with stochastic ensemble value expansion
J. Buckman, D. Hafner, G. Tucker, E. Brevdo, and H. Lee · 2018
Cited alongside, same era.
Deep reinforcement learning in a handful of trials using probabilistic dynamics models
K. Chua, R. Calandra, R. McAllister, and S. Levine · 2018
Cited alongside, same era.
Model-based value estimation for efficient model-free reinforcement learning
V. Feinberg, A. Wan, I. Stoica, M. I. Jordan, J. E. Gonzalez, and S. Levine · 2018
Cited alongside, same era.
Unsupervised meta-learning for reinforcement learning
A. Gupta, B. Eysenbach, C. Finn, and S. Levine · 2018
Cited alongside, same era.
Later among the works it cites.
Meta reinforcement learning with task embedding and shared policy
L. Lan, Z. Li, X. Guan, and P. Wang · 2019
Later among the works it cites.
A model-based approach for sample-efficient multi-task reinforcement learning
N. C. Landolfi, G. Thomas, and T. Ma · 2019
Later among the works it cites.
Guided meta-policy search
R. Mendonca, A. Gupta, R. Kralev, P. Abbeel, S. Levine, and C. Finn · 2019
Later among the works it cites.
Efficient off-policy meta-reinforcement learning via probabilistic context variables
K. Rakelly, A. Zhou, D. Quillen, C. Finn, and S. Levine · 2019
Later among the works it cites.
Exploring model-based planning with policy networks
T. Wang and J. Ba · 2019
Later among the works it cites.
Benchmarking model-based reinforcement learning
T. Wang, X. Bao, I. Clavera, J. Hoang, Y. Wen, E. Langlois, S. Zhang, G. Zhang, P. Abbeel, and J. Ba · 2019
Later among the works it cites.
Variational task embeddings for fast adapta-tion in deep reinforcement learning
L. Zintgraf, M. Igl, K. Shiarlis, A. Mahajan, K. Hofmann, and S. Whiteson · 2019
Later among the works it cites.
Reward-free exploration for reinforcement learning
C. Jin, A. Krishnamurthy, M. Simchowitz, and T. Yu · 2020
Closest in time.
Curriculum in gradient-based meta-reinforcement learning
B. Mehta, T. Deleu, S. C. Raparthy, C. J. Pal, and L. Paull · 2020
Closest in time.
R. Mendonca, X. Geng, C. Finn, and S. Levine · 2020
Closest in time.
A game theoretic framework for model based reinforcement learning
A. Rajeswaran, I. Mordatch, and V. Kumar · 2020
Closest in time.
Learning context-aware task reasoning for efficient meta-reinforcement learning
H. Wang, J. Zhou, and X. He · 2020
Closest in time.
Implicit function theorem — Wikipedia, the free encyclopedia, 2020
Wikipedia contributors · 2020
Closest in time.