Fetching the paper…
Reading the bibliography…
The transfer of knowledge from one policy to another is an important tool in Deep Reinforcement Learning.
Asynchronous methods for deep reinforcement learning
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., and Kavukcuoglu, K. (2016) · 1937
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
Ng, A. Y., Harada, D., and Russell, S. (1999) · 1999
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S., McAllester, D. A., Singh, S. P., and Mansour, Y. (2000) · 2000
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
Abbeel, P. and Ng, A. Y. (2004) · 2004
Earlier work this paper cites.
Divergence measures and message passing
Minka, T. et al. (2005) · 2005
Earlier work this paper cites.
Potential-based shaping in model-based reinforcement learning
Asmuth, J., Littman, M. L., and Zinkov, R. (2008) · 2008
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
Ross, S., Gordon, G., and Bagnell, D. (2011) · 2011
Earlier work this paper cites.
Dynamic potential-based reward shaping
Devlin, S. and Kudenko, D. (2012) · 2012
Earlier work this paper cites.
Distilling the knowledge in a neural network
Hinton, G., Vinyals, O., and Dean, J. (2015) · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al. (2015) · 2015
Earlier work this paper cites.
Actor-mimic: Deep multitask and transfer reinforcement learning
Parisotto, E., Ba, J. L., and Salakhutdinov, R. (2015) · 2015
Earlier work this paper cites.
Rusu, A. A., Colmenarejo, S. G., Gulcehre, C., Desjardins, G., Kirkpatrick, J., Pascanu, R., Mnih, V., Kavukcuoglu, K., and Hadsell, R. (2015) · 2015
Cited alongside, same era.
Mastering the game of go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al. (2016) · 2016
Cited alongside, same era.
Successor features for transfer in reinforcement learning
Barreto, A., Dabney, W., Munos, R., Hunt, J. J., Schaul, T., van Hasselt, H. P., and Silver, D. (2017) · 2017
Cited alongside, same era.
Gangwani, T. and Peng, J. (2017) · 2017
Cited alongside, same era.
Divide-and-conquer reinforcement learning
Ghosh, D., Singh, A., Rajeswaran, A., Kumar, V., and Levine, S. (2017) · 2017
Equivalence between policy gradients and soft q-learning
Schulman, J., Chen, X., and Abbeel, P. (2017) · 2017
Later among the works it cites.
Distral: Robust multitask reinforcement learning
Teh, Y., Bapst, V., Czarnecki, W. M., Quan, J., Kirkpatrick, J., Hadsell, R., Heess, N., and Pascanu, R. (2017) · 2017
Later among the works it cites.
Knowledge transfer for deep reinforcement learning with hierarchical experience replay
Yin, H. and Pan, S. J. (2017) · 2017
Later among the works it cites.
Zhang, Y., Xiang, T., Hospedales, T. M., and Lu, H. (2017) · 2017
Later among the works it cites.
Multi-task learning for continuous control
Arora, H., Kumar, R., Krone, J., and Li, C. (2018) · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Mobilenets: Efficient convolutional neural networks for mobile vision applications
Howard, A. G., Zhu, M., Chen, B., Kalenichenko, D., Wang, W., Weyand, T., Andreetto, M., and Adam, H. (2017) · 2017
Cited alongside, same era.
Reinforcement learning with unsupervised auxiliary tasks
Jaderberg, M., Mnih, V., Czarnecki, W. M., Schaul, T., Leibo, J. Z., Silver, D., and Kavukcuoglu, K. (2017) · 2017
Cited alongside, same era.
Collaborative deep reinforcement learning
Lin, K., Wang, S., and Zhou, J. (2017) · 2017
Cited alongside, same era.
Parallel wavenet: Fast high-fidelity speech synthesis
Oord, A. v. d., Li, Y., Babuschkin, I., Simonyan, K., Vinyals, O., Kavukcuoglu, K., Driessche, G. v. d., Lockhart, E., Cobo, L. C., Stimberg, F., et al. (2017) · 2017
Cited alongside, same era.
Curiosity-driven exploration by self-supervised prediction
Pathak, D., Agrawal, P., Efros, A. A., and Darrell, T. (2017) · 2017
Cited alongside, same era.
Evolution strategies as a scalable alternative to reinforcement learning
Salimans, T., Ho, J., Chen, X., Sidor, S., and Sutskever, I. (2017) · 2017
Cited alongside, same era.
Gradient estimation using stochastic computation graphs
Schulman, J., Heess, N., Weber, T., and Abbeel, P. (2015a)
Cited in the paper.
Mix&match-agent curricula for reinforcement learning
Czarnecki, W. M., Jayakumar, S. M., Jaderberg, M., Hasenclever, L., Teh, Y. W., Osindero, S., Heess, N., and Pascanu, R. (2018) · 2018
Later among the works it cites.
Scalable distributed deep-rl with importance weighted actor-learner architectures
Espeholt, L., Soyer, H., Munos, R., Simonyan, K., Mnih, V., Ward, T., Doron, Y., Firoiu, V., Harley, T., Dunning, I., Legg, S., and Kavukcuoglu, K. (2018) · 2018
Later among the works it cites.
Learning dexterous in-hand manipulation
OpenAI, Andrychowicz, M., Baker, B., Chociej, M., Jozefowicz, R., McGrew, B., Pachocki, J., Petron, A., Plappert, M., Powell, G., Ray, A., Schneider, J., Sidor, S., Tobin, J., Welinder, P., Weng, L., and Zaremba, W. (2018) · 2018
Later among the works it cites.
Model compression via distillation and quantization
Polino, A., Pascanu, R., and Alistarh, D. (2018) · 2018
Later among the works it cites.
Kickstarting deep reinforcement learning
Schmitt, S., Hudson, J. J., Zidek, A., Osindero, S., Doersch, C., Czarnecki, W. M., Leibo, J. Z., Kuttler, H., Zisserman, A., Simonyan, K., et al. (2018) · 2018
Later among the works it cites.