Fetching the paper…
Reading the bibliography…
Although reinforcement learning methods can achieve impressive results in simulation, the real world presents two major challenges: generating samples is exceedingly expensive, and unexpected perturbations or unseen situations cause proficient but specialized policies to fail at test time.
Adaptive control of linearizable systems
S. S. Sastry and A. Isidori · 1989
Earlier work this paper cites.
Learning a synaptic learning rule
Y. Bengio, S. Bengio, and J. Cloutier · 1990
Earlier work this paper cites.
Learning to generate artificial fovea trajectories for target detection
J. Schmidhuber and R. Huber · 1991
Earlier work this paper cites.
Meta-neural networks that learn by learning
D. K. Naik and R. Mammone · 1992
Earlier work this paper cites.
Learning to control fast-weight memories: An alternative to dynamic recurrent networks
J. Schmidhuber · 1992
Earlier work this paper cites.
Learning to learn: Introduction and overview
S. Thrun and L. Pratt · 1998
Earlier work this paper cites.
Meta-learning with backpropagation
A. S. Younger, S. Hochreiter, and P. R. Conwell · 2001
Earlier work this paper cites.
Reinforcement learning of motor skills with policy gradients
J. Peters and S. Schaal · 2008
Earlier work this paper cites.
Learning optimal adaptation strategies in unpredictable motor tasks
D. A. Braun, A. Aertsen, D. M. Wolpert, and C. Mehring · 2009
Earlier work this paper cites.
Gp-bayesfilters: Bayesian filtering using gaussian process prediction and observation models
J. Ko and D. Fox · 2009
Earlier work this paper cites.
Online parameter estimation and adaptive control of permanent-magnet synchronous machines
S. J. Underwood and I. Husain · 2010
Earlier work this paper cites.
Pilco: A model-based and data-efficient approach to policy search
M. Deisenroth and C. E. Rasmussen · 2011
Earlier work this paper cites.
Online movement adaptation based on previous sensor experiences
P. Pastor, L. Righetti, M. Kalakrishnan, and S. Schaal · 2011
Earlier work this paper cites.
Extensions of learning-based model predictive control for real-time application to a quadrotor helicopter
A. Aswani, P. Bouffard, and C. Tomlin · 2012
Earlier work this paper cites.
Online system identification and adaptive control for pem fuel cell maximum efficiency tracking
S. Kelouwani, K. Adegnon, K. Agbossou, and Y. Dube · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
E. Todorov, T. Erez, and Y. Tassa · 2012
Earlier work this paper cites.
Adaptive control
K. J. Åström and B. Wittenmark · 2013
Earlier work this paper cites.
A survey on policy search for robotics
M. P. Deisenroth, G. Neumann, J. Peters, et al · 2013
Earlier work this paper cites.
Guided policy search
S. Levine and V. Koltun · 2013
Earlier work this paper cites.
Adaptive model predictive control for constrained linear systems
M. Tanaskovic, L. Fagiano, R. Smith, P. Goulart, and M. Morari · 2013
Earlier work this paper cites.
Optimization of perturbative pv mppt methods through online system identification
P. Manganiello, M. Ricco, G. Petrone, E. Monmasson, and G. Spagnuolo · 2014
Cited alongside, same era.
One-shot learning of manipulation skills with online dynamics adaptation and neural network priors
J. Fu, S. Levine, and P. Abbeel · 2015
Cited alongside, same era.
Human-level concept learning through probabilistic program induction
B. M. Lake, R. Salakhutdinov, and J. B. Tenenbaum · 2015
Cited alongside, same era.
Deepmpc: Learning deep latent features for model predictive control
I. Lenz, R. A. Knepper, and A. Saxena · 2015
Cited alongside, same era.
Continuous control with deep reinforcement learning
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra · 2015
Cited alongside, same era.
Learning to reinforcement learn
J. X. Wang, Z. Kurth-Nelson, D. Tirumala, H. Soyer, J. Z. Leibo, R. Munos, C. Blundell, D. Kumaran, and M. Botvinick · 2016
Later among the works it cites.
Continuous adaptation via meta-learning in nonstationary and competitive environments
M. Al-Shedivat, T. Bansal, Y. Burda, I. Sutskever, I. Mordatch, and P. Abbeel · 2017
Later among the works it cites.
Model-based policy search for automatic tuning of multivariate PID controllers
A. Doerr, D. Nguyen-Tuong, A. Marco, S. Schaal, and S. Trimpe · 2017
Later among the works it cites.
C. Finn and S. Levine · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al · 2015
Cited alongside, same era.
Online representation learning in recurrent neural language models
M. Rei · 2015
Cited alongside, same era.
Trust region policy optimization
J. Schulman, S. Levine, P. Moritz, M. I. Jordan, and P. Abbeel · 2015
Cited alongside, same era.
Model predictive path integral control using covariance variable importance sampling
G. Williams, A. Aldrich, and E. Theodorou · 2015
Cited alongside, same era.
Learning to learn by gradient descent by gradient descent
M. Andrychowicz, M. Denil, S. G. Colmenarejo, M. W. Hoffman, D. Pfau, T. Schaul, and N. de Freitas · 2016
Cited alongside, same era.
Designing neural network architectures using reinforcement learning
B. Baker, O. Gupta, N. Naik, and R. Raskar · 2016
Cited alongside, same era.
Rl$ˆ2$: Fast reinforcement learning via slow reinforcement learning
Y. Duan, J. Schulman, X. Chen, P. L. Bartlett, I. Sutskever, and P. Abbeel · 2016
Cited alongside, same era.
C. Finn, P. Abbeel, and S. Levine · 2017
Later among the works it cites.
Bayesian recurrent neural networks
M. Fortunato, C. Blundell, and O. Vinyals · 2017
Later among the works it cites.
Dynamic evaluation of neural sequence models
B. Krause, E. Kahembwe, I. Murray, and S. Renals · 2017
Later among the works it cites.
A simple neural attentive meta-learner
N. Mishra, M. Rohaninejad, X. Chen, and P. Abbeel · 2017
Later among the works it cites.
T. Munkhdalai and H. Yu · 2017
Later among the works it cites.
Learning rapid-temporal adaptations
T. Munkhdalai, X. Yuan, S. Mehri, T. Wang, and A. Trischler · 2017
Later among the works it cites.
Learning feedback terms for reactive planning and control
A. Rai, G. Sutanto, S. Schaal, and F. Meier · 2017
Later among the works it cites.
Mastering the game of go without human knowledge
D. Silver, J. Schrittwieser, K. Simonyan, I. Antonoglou, A. Huang, A. Guez, T. Hubert, L. Baker, M. Lai, A. Bolton, et al · 2017
Later among the works it cites.
Learning to learn: Meta-critic networks for sample efficient learning
F. Sung, L. Zhang, T. Xiang, T. Hospedales, and Y. Yang · 2017
Later among the works it cites.
Structure learning in motor control: A deep reinforcement learning model
A. Weinstein and M. Botvinick · 2017
Later among the works it cites.
Information theoretic mpc for model-based reinforcement learning
G. Williams, N. Wagener, B. Goldfain, P. Drews, J. M. Rehg, B. Boots, and E. A. Theodorou · 2017
Later among the works it cites.
Deep reinforcement learning in a handful of trials using probabilistic dynamics models
K. Chua, R. Calandra, R. McAllister, and S. Levine · 2018
Closest in time.
Model-ensemble trust-region policy optimization
T. Kurutach, I. Clavera, Y. Duan, A. Tamar, and P. Abbeel · 2018
Closest in time.
Optimization as a model for few-shot learning
S. Ravi and H. Larochelle · 2018
Closest in time.
Meta reinforcement learning with latent variable gaussian processes
S. Sæmundsson, K. Hofmann, and M. P. Deisenroth · 2018
Closest in time.