Fetching the paper…
Reading the bibliography…
Being able to seamlessly generalize across different tasks is fundamental for robots to act in our world.
Evolutionary principles in self-referential learning, or on learning how to learn: the meta-meta-… hook. Institut für Informatik, Technische Universität München, 1987
Schmidhuber, J · 1987
Earlier work this paper cites.
Learning a synaptic learning rule
Bengio, Y. and Bengio, S · 1990
Earlier work this paper cites.
A comparison of direct and model-based reinforcement learning
Atkeson, C. G. and Santamaria, J. C · 1997
Earlier work this paper cites.
Neural fitted q iteration–first experiences with a data efficient neural reinforcement learning method
Riedmiller, M · 2005
Earlier work this paper cites.
Reinforcement learning by value gradients
Fairbank, M · 2008
Earlier work this paper cites.
Pilco: A model-based and data-efficient approach to policy search
Deisenroth, M. and Rasmussen, C. E · 2011
Earlier work this paper cites.
Learning to learn
Thrun, S. and Pratt, L · 2012
Earlier work this paper cites.
Deterministic policy gradient algorithms
Silver, D., Lever, G., Heess, N., Degris, T., Wierstra, D., and Riedmiller, M · 2014
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D · 2015
Earlier work this paper cites.
Gradient-based hyperparameter optimization through reversible learning
Maclaurin, D., Duvenaud, D., and Adams, R · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Earlier work this paper cites.
Universal value function approximators
Schaul, T., Horgan, D., Gregor, K., and Silver, D · 2015
Earlier work this paper cites.
Learning to learn by gradient descent by gradient descent
Andrychowicz, M., Denil, M., Colmenarejo, S. G., Hoffman, M. W., Pfau, D., Schaul, T., and de Freitas, N · 2016
Earlier work this paper cites.
Rl 2 : Fast reinforcement learning via slow reinforcement learning
Duan, Y., Schulman, J., Chen, X., Bartlett, P. L., Sutskever, I., and Abbeel, P · 2016
Cited alongside, same era.
Li, K. and Malik, J · 2016
Cited alongside, same era.
Andrychowicz, M., Wolski, F., Ray, A., Schneider, J., Fong, R., Welinder, P., McGrew, B., Tobin, J., Abbeel, P., and Zaremba, W · 2017
Cited alongside, same era.
Model-agnostic meta-learning for fast adaptation of deep networks
Finn, C., Abbeel, P., and Levine, S · 2017
Cited alongside, same era.
Forward and reverse gradient-based hyperparameter optimization
Franceschi, L., Donini, M., Frasconi, P., and Pontil, M · 2017
Cited alongside, same era.
Real2sim transfer using differentiable physics
Heiden, E., Millard, D., and Sukhatme, G · 2019
Later among the works it cites.
Mendonca, R., Gupta, A., Kralev, R., Abbeel, P., Levine, S., and Finn, C · 2019
Later among the works it cites.
Investigating generalisation in continuous deep reinforcement learning
Zhao, C., Sigaud, O., Stulp, F., and Hospedales, T. M · 2019
Later among the works it cites.
Reward shaping via meta-learning
Zou, H., Ren, T., Yan, D., Su, H., and Zhu, J · 2019
Later among the works it cites.
Curious ilqr: Resolving uncertainty in model-based rl
Bechtle, S., Lin, Y., Rai, A., Righetti, L., and Meier, F · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learning to learn: Meta-critic networks for sample efficient learning
Sung, F., Zhang, L., Xiang, T., Hospedales, T., and Yang, Y · 2017
Cited alongside, same era.
Deep reinforcement learning in a handful of trials using probabilistic dynamics models
Chua, K., Calandra, R., McAllister, R., and Levine, S · 2018
Cited alongside, same era.
Meta-reinforcement learning of structured exploration strategies
Gupta, A., Mendonca, R., Liu, Y., Abbeel, P., and Levine, S · 2018
Cited alongside, same era.
Evolved policy gradients
Houthooft, R., Chen, Y., Isola, P., Stadie, B. C., Wolski, F., Ho, J., and Abbeel, P · 2018
Cited alongside, same era.
Online learning of a memory for learning rates
Meier, F., Kappler, D., and Schaal, S · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G · 2018
Cited alongside, same era.
One-shot imitation from observing humans via domain-adaptive meta-learning
Yu, T., Finn, C., Xie, A., Dasari, S., Zhang, T., Abbeel, P., and Levine, S · 2018
Cited alongside, same era.
Hospedales, T., Antoniou, A., Micaelli, P., and Storkey, A · 2020
Later among the works it cites.
Encoding physical constraints in differentiable newton-euler algorithm
Sutanto, G., Wang, A., Lin, Y., Mukadam, M., Sukhatme, G., Rai, A., and Meier, F · 2020
Later among the works it cites.
Meta-gradient reinforcement learning with an objective discovered online
Xu, Z., van Hasselt, H., Hessel, M., Oh, J., Singh, S., and Silver, D · 2020
Later among the works it cites.
Online meta-critic learning for off-policy actor-critic methods
Zhou, W., Li, Y., Yang, Y., Wang, H., and Hospedales, T. M · 2020
Later among the works it cites.
Meta learning via learned loss
Bechtle, S., Molchanov, A., Chebotar, Y., Grefenstette, E., Righetti, L., Sukhatme, G., and Meier, F · 2021
Later among the works it cites.
Neuralsim: Augmenting differentiable simulators with neural networks
Heiden, E., Millard, D., Coumans, E., Sheng, Y., and Sukhatme, G. S · 2021
Later among the works it cites.
A survey of generalisation in deep reinforcement learning
Kirk, R., Zhang, A., Grefenstette, E., and Rocktäschel, T · 2021
Later among the works it cites.