Fetching the paper…
Reading the bibliography…
Deep reinforcement learning algorithms require large amounts of experience to learn an individual task.
Bayesian model-agnostic meta-learning
Yoon, J., Kim, T., Dia, O., Kim, S., Bengio, Y., and Ahn, S · 1903
Earlier work this paper cites.
Evolutionary principles in self-referential learning, or on learning how to learn: the meta-meta-… hook
Schmidhuber, J · 1987
Earlier work this paper cites.
Learning a synaptic learning rule
Bengio, Y., Bengio, S., and Cloutier, J · 1990
Earlier work this paper cites.
Planning and acting in partially observable stochastic domains
Kaelbling, L. P., Littman, M., and Cassandra, A · 1998
Earlier work this paper cites.
Learning to learn
Thrun, S. and Pratt, L · 1998
Earlier work this paper cites.
A Bayesian framework for concept learning
Tenenbaum, J. B · 1999
Earlier work this paper cites.
A bayesian framework for reinforcement learning
Strens, M · 2000
Earlier work this paper cites.
A bayesian approach to unsupervised one-shot learning of object categories
Fei-Fei, L. et al · 2003
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Todorov, E., Erez, T., and Tassa, Y · 2012
Earlier work this paper cites.
(more) efficient reinforcement learning via posterior sampling
Osband, I., Van Roy, B., and Russo, D · 2013
Earlier work this paper cites.
Auto-encoding variational bayes
Kingma, D. P. and Welling, M · 2014
Earlier work this paper cites.
Stochastic backpropagation and approximate inference in deep generative models
Rezende, D. J., Mohamed, S., and Wierstra, D · 2014
Earlier work this paper cites.
Deep recurrent q-learning for partially observable mdps
Hausknecht, M. and Stone, P · 2015
Earlier work this paper cites.
Memory-based control with recurrent neural networks
Heess, N., Hunt, J. J., Lillicrap, T. P., and Silver, D · 2015
Earlier work this paper cites.
Deep variational information bottleneck
Alemi, A. A., Fischer, I., Dillon, J. V., and Murphy, K · 2016
Cited alongside, same era.
Rl 2 : Fast reinforcement learning via slow reinforcement learning
Duan, Y., Schulman, J., Chen, X., Bartlett, P. L., Sutskever, I., and Abbeel, P · 2016
Cited alongside, same era.
Deep exploration via bootstrapped dqn
Osband, I., Blundell, C., Pritzel, A., and Van Roy, B · 2016
Cited alongside, same era.
Meta-learning with memory-augmented neural networks
Santoro, A., Bartunov, S., Botvinick, M., Wierstra, D., and Lillicrap, T · 2016
Cited alongside, same era.
Matching networks for one shot learning
Vinyals, O., Blundell, C., Lillicrap, T., Wierstra, D., et al · 2016
Cited alongside, same era.
Meta-reinforcement learning of structured exploration strategies
Gupta, A., Mendonca, R., Liu, Y., Abbeel, P., and Levine, S · 2018
Later among the works it cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S · 2018
Later among the works it cites.
Learning an embedding space for transferable robot skills
Hausman, K., Springenberg, J. T., Wang, Z., Heess, N., and Riedmiller, M · 2018
Later among the works it cites.
Evolved policy gradients
Houthooft, R., Chen, R. Y., Isola, P., Stadie, B. C., Wolski, F., Ho, J., and Abbeel, P · 2018
Later among the works it cites.
Deep variational reinforcement learning for pomdps
Igl, M., Zintgraf, L., Le, T. A., Wood, F., and Whiteson, S · 2018
Later among the works it cites.
Task-embedded control networks for few-shot imitation learning
James, S., Bloesch, M., and Davison, A. J · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Wang, J. X., Kurth-Nelson, Z., Tirumala, D., Soyer, H., Leibo, J. Z., Munos, R., Blundell, C., Kumaran, D., and Botvinick, M · 2016
Cited alongside, same era.
One-shot imitation learning
Duan, Y., Andrychowicz, M., Stadie, B., Ho, O. J., Schneider, J., Sutskever, I., Abbeel, P., and Zaremba, W · 2017
Cited alongside, same era.
Model-agnostic meta-learning for fast adaptation of deep networks
Finn, C., Abbeel, P., and Levine, S · 2017
Cited alongside, same era.
Optimization as a model for few-shot learning
Ravi, S. and Larochelle, H · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P. D., Radford, A. R., and Klimov, O · 2017
Cited alongside, same era.
Prototypical networks for few-shot learning
Snell, J., Swersky, K., and Zemel, R · 2017
Cited alongside, same era.
Learning to learn: Meta-critic networks for sample efficient learning
Sung, F., Zhang, L., Xiang, T., Hospedales, T., and Yang, Y · 2017
Cited alongside, same era.
Later among the works it cites.
A simple neural attentive meta-learner
Mishra, N., Rohaninejad, M., Chen, X., and Abbeel, P · 2018
Later among the works it cites.
Tadam: Task dependent adaptive metric for improved few-shot learning
Oreshkin, B. N., Lacoste, A., and Rodriguez, P · 2018
Later among the works it cites.
Promp: Proximal meta-policy search
Rothfuss, J., Lee, D., Clavera, I., Asfour, T., and Abbeel, P · 2018
Later among the works it cites.
Meta reinforcement learning with latent variable gaussian processes
Sæmundsson, S., Hofmann, K., and Deisenroth, M. P · 2018
Later among the works it cites.
Some considerations on learning to explore via meta-reinforcement learning
Stadie, B. C., Yang, G., Houthooft, R., Chen, X., Duan, Y., Wu, Y., Abbeel, P., and Sutskever, I · 2018
Later among the works it cites.
Meta-learning probabilistic inference for prediction
Gordon, J., Bronskill, J., Bauer, M., Nowozin, S., and Turner, R · 2019
Closest in time.
Learning to adapt in dynamic, real-world environments through meta-reinforcement learning
Nagabandi, A., Clavera, I., Liu, S., Fearing, R. S., Abbeel, P., Levine, S., and Finn, C · 2019
Closest in time.
Meta-learning with latent embedding optimization
Rusu, A. A., Rao, D., Sygnowski, J., Vinyals, O., Pascanu, R., Osindero, S., and Hadsell, R · 2019
Closest in time.