Fetching the paper…
Reading the bibliography…
Policy gradient methods have shown success in learning control policies for high-dimensional dynamical systems.
Quasi-martingales
Fisk, D. (1965) · 1965
Earlier work this paper cites.
Convergence of probability measures
Billingsley, P. (1968) · 1968
Earlier work this paper cites.
Catastrophic interference in connectionist networks: The sequential learning problem
McCloskey, M. and Cohen, N. J. (1989) · 1989
Earlier work this paper cites.
Optimization problems with perturbations: A guided tour
Bonnans, J. F. and Shapiro, A. (1998) · 1998
Earlier work this paper cites.
Asymptotic statistics
Van der Vaart, A. (2000) · 2000
Earlier work this paper cites.
A natural policy gradient
Kakade, S. M. (2002) · 2002
Earlier work this paper cites.
Recovery of exact sparse representations in the presence of bounded noise
Fuchs, J.-J. (2005) · 2005
Earlier work this paper cites.
The strong law of large numbers for L-statistics with dependent data
Baklanov, E. A. (2006) · 2006
Earlier work this paper cites.
On-line learning and stochastic approximations
Bottou, L. (2009) · 2009
Earlier work this paper cites.
MuJoCo: A physics engine for model-based control
Todorov, E., Erez, T., and Tassa, Y. (2012) · 2012
Earlier work this paper cites.
ELLA: An efficient lifelong learning algorithm
Ruvolo, P. and Eaton, E. (2013) · 2013
Earlier work this paper cites.
Online multi-task learning for policy gradient methods
Bou Ammar, H., Eaton, E., Ruvolo, P., and Taylor, M. (2014) · 2014
Earlier work this paper cites.
Autonomous cross-domain knowledge transfer in lifelong policy gradient reinforcement learning
Bou Ammar, H., Eaton, E., Luna, J. M., and Ruvolo, P. (2015) · 2015
Earlier work this paper cites.
Trust region policy optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M., and Moritz, P. (2015) · 2015
Earlier work this paper cites.
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W. (2016) · 2016
Earlier work this paper cites.
RL 2 : Fast reinforcement learning via slow reinforcement learning
Duan, Y., Schulman, J., Chen, X., Bartlett, P. L., Sutskever, I., and Abbeel, P. (2016) · 2016
Cited alongside, same era.
Using task features for zero-shot knowledge transfer in lifelong learning
Isele, D., Rostami, M., and Eaton, E. (2016) · 2016
Cited alongside, same era.
Continuous control with deep reinforcement learning
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D. (2016) · 2016
Cited alongside, same era.
Actor-mimic: Deep multitask and transfer reinforcement learning
Parisotto, E., Ba, J. L., and Salakhutdinov, R. (2016) · 2016
Cited alongside, same era.
Rusu, A. A., Rabinowitz, N. C., Desjardins, G., Soyer, H., Kirkpatrick, J., Kavukcuoglu, K., Pascanu, R., and Hadsell, R. (2016) · 2016
Cited alongside, same era.
Tensor based knowledge transfer across skill categories for robot control
Zhao, C., Hospedales, T. M., Stulp, F., and Sigaud, O. (2017) · 2017
Later among the works it cites.
Meta-reinforcement learning of structured exploration strategies
Gupta, A., Mendonca, R., Liu, Y., Abbeel, P., and Levine, S. (2018) · 2018
Later among the works it cites.
Note on the quadratic penalties in elastic weight consolidation
Huszár, F. (2018) · 2018
Later among the works it cites.
Selective experience replay for lifelong learning
Isele, D. and Cosgun, A. (2018) · 2018
Later among the works it cites.
Variational continual learning
Nguyen, C. V., Li, Y., Bui, T. D., and Turner, R. E. (2018) · 2018
Later among the works it cites.
Online structured Laplace approximations for overcoming catastrophic forgetting
Ritter, H., Botev, A., and Barber, D. (2018) · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Model-agnostic meta-learning for fast adaptation of deep networks
Finn, C., Abbeel, P., and Levine, S. (2017) · 2017
Cited alongside, same era.
Benchmark environments for multitask learning in continuous domains
Henderson, P., Chang, W.-D., Shkurti, F., Hansen, J., Meger, D., and Dudek, G. (2017) · 2017
Cited alongside, same era.
Overcoming catastrophic forgetting in neural networks
Kirkpatrick, J., Pascanu, R., Rabinowitz, N., Veness, J., Desjardins, G., Rusu, A. A., Milan, K., Quan, J., Ramalho, T., Grabska-Barwinska, A., et al. (2017) · 2017
Cited alongside, same era.
Learning without forgetting
Li, Z. and Hoiem, D. (2017) · 2017
Cited alongside, same era.
Towards generalization and simplicity in continuous control
Rajeswaran, A., Lowrey, K., Todorov, E. V., and Kakade, S. M. (2017) · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O. (2017) · 2017
Cited alongside, same era.
Distral: Robust multitask reinforcement learning
Teh, Y., Bapst, V., Czarnecki, W. M., Quan, J., Kirkpatrick, J., Hadsell, R., Heess, N., and Pascanu, R. (2017) · 2017
Cited alongside, same era.
Progress & compress: A scalable framework for continual learning
Schwarz, J., Czarnecki, W., Luketina, J., Grabska-Barwinska, A., Teh, Y. W., Pascanu, R., and Hadsell, R. (2018) · 2018
Later among the works it cites.
Learning to adapt in dynamic, real-world environments through meta-reinforcement learning
Clavera, I., Nagabandi, A., Liu, S., Fearing, R. S., Abbeel, P., Levine, S., and Finn, C. (2019) · 2019
Later among the works it cites.
A meta-MDP approach to exploration for lifelong reinforcement learning
Garcia, F. and Thomas, P. S. (2019) · 2019
Later among the works it cites.
Deep online learning via meta-learning: Continual adaptation for model-based RL
Nagabandi, A., Finn, C., and Levine, S. (2019) · 2019
Later among the works it cites.
Experience replay for continual learning
Rolnick, D., Ahuja, A., Schwarz, J., Lillicrap, T., and Wayne, G. (2019) · 2019
Later among the works it cites.
Meta-World: A benchmark and evaluation for multi-task and meta reinforcement learning
Yu, T., Quillen, D., He, Z., Julian, R., Hausman, K., Finn, C., and Levine, S. (2019) · 2019
Later among the works it cites.
Uncertainty-guided continual learning with Bayesian neural networks
Ebrahimi, S., Elhoseiny, M., Darrell, T., and Rohrbach, M. (2020) · 2020
Closest in time.
Functional regularisation for continual learning with Gaussian processes
Titsias, M. K., Schwarz, J., de G. Matthews, A. G., Pascanu, R., and Teh, Y. W. (2020) · 2020
Closest in time.