Fetching the paper…
Reading the bibliography…
The goal of meta-reinforcement learning (meta-RL) is to build agents that can quickly learn new tasks by leveraging prior experience on related tasks.
Watch, try, learn: Meta-learning from demonstrations and reward
Zhou, A., Jang, E., Kappler, D., Herzog, A., Khansari, M., Wohlhart, P., Bai, Y., Kalakrishnan, M., Levine, S., and Finn, C · 1906
Earlier work this paper cites.
Environment probing interaction policies
Zhou, W., Pinto, L., and Gupta, A · 1907
Earlier work this paper cites.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
Thompson, W. R · 1933
Earlier work this paper cites.
Evolutionary principles in self-referential learning, or on learning how to learn: the meta-meta-… hook
Schmidhuber, J · 1987
Earlier work this paper cites.
Learning a synaptic learning rule
Bengio, Y., Bengio, S., and Cloutier, J · 1991
Earlier work this paper cites.
On the optimization of a synaptic learning rule
Bengio, S., Bengio, Y., Cloutier, J., and Gecsei, J · 1992
Earlier work this paper cites.
Meta-neural networks that learn by learning
Naik, D. K. and Mammone, R. J · 1992
Earlier work this paper cites.
Planning and acting in partially observable stochastic domains
Kaelbling, L. P., Littman, M. L., and Cassandra, A. R · 1998
Earlier work this paper cites.
Learning to learn using gradient descent
Hochreiter, S., Younger, A. S., and Conwell, P. R · 2001
Earlier work this paper cites.
The IM algorithm: a variational approach to information maximization
Barber, D. and Agakov, F. V · 2003
Earlier work this paper cites.
An imitation learning approach for cache replacement
Liu, E. Z., Hashemi, M., Swersky, K., Ranganathan, P., and Ahn, J · 2006
Earlier work this paper cites.
Learning abstract models for strategic exploration and fast reward transfer
Liu, E. Z., Keramati, R., Seshadri, S., Guu, K., Pasupat, P., Brunskill, E., and Liang, P · 2007
Earlier work this paper cites.
Visualizing data using t-SNE
van der Maaten, L. and Hinton, G · 2008
Earlier work this paper cites.
Learning to learn
Thrun, S. and Pratt, L · 2012
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. and Ba, J · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Earlier work this paper cites.
Deep variational information bottleneck
Alemi, A. A., Fischer, I., Dillon, J. V., and Murphy, K · 2016
Earlier work this paper cites.
Learning to learn by gradient descent by gradient descent
Andrychowicz, M., Denil, M., Gomez, S., Hoffman, M. W., Pfau, D., Schaul, T., Shillingford, B., and Freitas, N. D · 2016
Earlier work this paper cites.
Unifying count-based exploration and intrinsic motivation
Bellemare, M., Srinivasan, S., Ostrovski, G., Schaul, T., Saxton, D., and Munos, R · 2016
Cited alongside, same era.
RL 2 : Fast reinforcement learning via slow reinforcement learning
Duan, Y., Schulman, J., Chen, X., Bartlett, P. L., Sutskever, I., and Abbeel, P · 2016
Cited alongside, same era.
Gregor, K., Rezende, D. J., and Wierstra, D · 2016
Cited alongside, same era.
Vime: Variational information maximizing exploration
Houthooft, R., Chen, X., Duan, Y., Schulman, J., Turck, F. D., and Abbeel, P · 2016
Cited alongside, same era.
One-shot learning with memory-augmented neural networks
Santoro, A., Bartunov, S., Botvinick, M., Wierstra, D., and Lillicrap, T · 2016
Cited alongside, same era.
Meta reinforcement learning with latent variable gaussian processes
Sæmundsson, S., Hofmann, K., and Deisenroth, M. P · 2018
Later among the works it cites.
The importance of sampling inmeta-reinforcement learning
Stadie, B., Yang, G., Houthooft, R., Chen, P., Duan, Y., Wu, Y., Abbeel, P., and Sutskever, I · 2018
Later among the works it cites.
Unsupervised control through non-parametric discriminative rewards
Warde-Farley, D., de Wiele, T. V., Kulkarni, T., Ionescu, C., Hansen, S., and Mnih, V · 2018
Later among the works it cites.
Learning to generalize from sparse and underspecified rewards
Agarwal, R., Liang, C., Schuurmans, D., and Norouzi, M · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deep reinforcement learning with double Q-learning
van Hasselt, H., Guez, A., and Silver, D · 2016
Cited alongside, same era.
Model-agnostic meta-learning for fast adaptation of deep networks
Finn, C., Abbeel, P., and Levine, S · 2017
Cited alongside, same era.
A simple neural attentive meta-learner
Mishra, N., Rohaninejad, M., Chen, X., and Abbeel, P · 2017
Cited alongside, same era.
Automatic differentiation in pytorch, 2017
Paszke, A., Gross, S., Chintala, S., Chanan, G., Yang, E., DeVito, Z., Lin, Z., Desmaison, A., Antiga, L., and Lerer, A · 2017
Cited alongside, same era.
Curiosity-driven exploration by self-supervised prediction
Pathak, D., Agrawal, P., Efros, A. A., and Darrell, T · 2017
Cited alongside, same era.
A tutorial on thompson sampling
Russo, D., Roy, B. V., Kazerouni, A., Osband, I., and Wen, Z · 2017
Cited alongside, same era.
Exploration by random network distillation
Burda, Y., Edwards, H., Storkey, A., and Klimov, O · 2018
Cited alongside, same era.
Fakoor, R., Chaudhari, P., Soatto, S., and Smola, A. J · 2019
Later among the works it cites.
Mame: Model-agnostic meta-exploration
Gurumurthy, S., Kumar, S., and Sycara, K · 2019
Later among the works it cites.
Meta reinforcement learning as task inference
Humplik, J., Galashov, A., Hasenclever, L., Ortega, P. A., Teh, Y. W., and Heess, N · 2019
Later among the works it cites.
Recurrent experience replay in distributed reinforcement learning
Kapturowski, S., Ostrovski, G., Quan, J., Munos, R., and Dabney, W · 2019
Later among the works it cites.
A unified bellman optimality principle combining reward maximization and empowerment
Leibfried, F., Pascual-Diaz, S., and Grau-Moya, J · 2019
Later among the works it cites.
Guided meta-policy search
Mendonca, R., Gupta, A., Kralev, R., Abbeel, P., Levine, S., and Finn, C · 2019
Later among the works it cites.
Efficient off-policy meta-reinforcement learning via probabilistic context variables
Rakelly, K., Zhou, A., Quillen, D., Finn, C., and Levine, S · 2019
Later among the works it cites.
Norml: No-reward meta learning
Yang, Y., Caluwaerts, K., Iscen, A., Tan, J., and Finn, C · 2019
Later among the works it cites.
Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning
Yu, T., Quillen, D., He, Z., Julian, R., Hausman, K., Finn, C., and Levine, S · 2019
Later among the works it cites.
Varibad: A very good method for bayes-adaptive deep RL via meta-learning
Zintgraf, L., Shiarlis, K., Igl, M., Schulze, S., Gal, Y., Hofmann, K., and Whiteson, S · 2019
Later among the works it cites.
Offline meta reinforcement learning
Dorfman, R. and Tamar, A · 2020
Closest in time.
Meta-model-based meta-policy optimization
Hiraoka, T., Imagawa, T., Tangkaratt, V., Osa, T., Onishi, T., and Tsuruoka, Y · 2020
Closest in time.
Kamienny, P.-A., Pirotta, M., Lazaric, A., Lavril, T., Usunier, N., and Denoyer, L · 2020
Closest in time.
Learn to effectively explore in context-based meta-RL
Zhang, J., Wang, J., Hu, H., Chen, Y., Fan, C., and Zhang, C · 2020
Closest in time.