Fetching the paper…
Reading the bibliography…
Standard RL algorithms assume fixed environment dynamics and require a significant amount of interaction to adapt to new environments.
Off-policy temporal-difference learning with function approximation
Precup, D., Sutton, R. S., and Dasgupta, S · 2001
Earlier work this paper cites.
Visualizing data using t-sne
van der Maaten, L. and Hinton, G. E · 2008
Earlier work this paper cites.
Transfer learning for reinforcement learning domains: A survey
Taylor, M. E. and Stone, P · 2009
Earlier work this paper cites.
Horde: a scalable real-time architecture for learning knowledge from unsupervised sensorimotor interaction
Sutton, R. S., Modayil, J., Delp, M., Degris, T., Pilarski, P. M., White, A., and Precup, D · 2011
Earlier work this paper cites.
Protecting against evaluation overfitting in empirical reinforcement learning
Whiteson, S., Tanner, B., Taylor, M. E., and Stone, P · 2011
Earlier work this paper cites.
Reinforcement learning transfer via sparse coding
Ammar, H. B., Tuyls, K., Taylor, M. E., Driessens, K., and Weiss, G · 2012
Earlier work this paper cites.
Da Silva, B., Konidaris, G., and Barto, A · 2012
Earlier work this paper cites.
Transfer in reinforcement learning: a framework and a survey
Lazaric, A · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Todorov, E., Erez, T., and Tassa, Y · 2012
Earlier work this paper cites.
Scaling life-long off-policy learning
White, A., Modayil, J., and Sutton, R. S · 2012
Earlier work this paper cites.
Hidden parameter markov decision processes: A semiparametric regression approach for discovering latent task parametrizations
Doshi-Velez, F. and Konidaris, G · 2013
Earlier work this paper cites.
Auto-encoding variational bayes
Kingma, D. P. and Welling, M · 2013
Earlier work this paper cites.
An automated measure of mdp similarity for transfer in reinforcement learning
Ammar, H. B., Eaton, E., Taylor, M. E., Mocanu, D. C., Driessens, K., Weiss, G., and Tuyls, K · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Robots that can adapt like animals
Cully, A., Clune, J., Tarapore, D., and Mouret, J.-B · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Earlier work this paper cites.
Actor-mimic: Deep multitask and transfer reinforcement learning
Parisotto, E., Ba, J. L., and Salakhutdinov, R · 2015
Earlier work this paper cites.
Universal value function approximators
Schaul, T., Horgan, D., Gregor, K., and Silver, D · 2015
Earlier work this paper cites.
Learning shared representations in multi-task reinforcement learning
Borsa, D., Graepel, T., and Shawe-Taylor, J · 2016
Earlier work this paper cites.
Learning modular neural network policies for multi-task and multi-robot transfer
Devin, C., Gupta, A., Darrell, T., Abbeel, P., and Levine, S · 2016
Earlier work this paper cites.
Rl 2 : Fast reinforcement learning via slow reinforcement learning
Duan, Y., Schulman, J., Chen, X., Bartlett, P. L., Sutskever, I., and Abbeel, P · 2016
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al · 2016
Earlier work this paper cites.
Learning to reinforcement learn
Wang, J. X., Kurth-Nelson, Z., Soyer, H., Leibo, J. Z., Tirumala, D., Munos, R., Blundell, C., Kumaran, D., and Botvinick, M. M · 2016
Earlier work this paper cites.
Modular multitask reinforcement learning with policy sketches
Andreas, J., Klein, D., and Levine, S · 2017
Cited alongside, same era.
Successor features for transfer in reinforcement learning
Barreto, A., Dabney, W., Munos, R., Hunt, J. J., Schaul, T., van Hasselt, H. P., and Silver, D · 2017
Cited alongside, same era.
One-shot imitation learning
Duan, Y., Andrychowicz, M., Stadie, B. C., Ho, J., Schneider, J., Sutskever, I., Abbeel, P., and Zaremba, W · 2017
Cited alongside, same era.
Model-agnostic meta-learning for fast adaptation of deep networks
Finn, C., Abbeel, P., and Levine, S · 2017
Cited alongside, same era.
Learning invariant feature spaces to transfer skills with reinforcement learning
Gupta, A., Devin, C., Liu, Y., Abbeel, P., and Levine, S · 2017
Cited alongside, same era.
Darla: Improving zero-shot transfer in reinforcement learning
Houthooft, R., Chen, R. Y., Isola, P., Stadie, B. C., Wolski, F., Ho, J., and Abbeel, P · 2018
Later among the works it cites.
Learning to adapt in dynamic, real-world environments through meta-reinforcement learning
Nagabandi, A., Clavera, I., Liu, S., Fearing, R. S., Abbeel, P., Levine, S., and Finn, C · 2018
Later among the works it cites.
Fingerprint policy optimisation for robust reinforcement learning
Paul, S., Osborne, M. A., and Whiteson, S · 2018
Later among the works it cites.
Efficient transfer learning and online adaptation with latent variable models for continuous control
Perez, C. F., Such, F. P., and Karaletsos, T · 2018
Later among the works it cites.
Meta reinforcement learning with latent variable gaussian processes
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Higgins, I., Pal, A., Rusu, A. A., Matthey, L., Burgess, C., Pritzel, A., Botvinick, M. M., Blundell, C., and Lerchner, A · 2017
Cited alongside, same era.
Robust and efficient transfer learning with hidden parameter markov decision processes
Killian, T. W., Konidaris, G., and Doshi-Velez, F · 2017
Cited alongside, same era.
Zero-shot task generalization with multi-task deep reinforcement learning
Oh, J., Singh, S., Lee, H., and Kohli, P · 2017
Cited alongside, same era.
Robust adversarial reinforcement learning
Pinto, L., Davidson, J., Sukthankar, R., and Gupta, A · 2017
Cited alongside, same era.
Towards generalization and simplicity in continuous control
Rajeswaran, A., Lowrey, K., Todorov, E. V., and Kakade, S. M · 2017
Cited alongside, same era.
Sahni, H., Kumar, S., Tejani, F., and Isbell, C · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Cited alongside, same era.
Sæmundsson, S., Hofmann, K., and Deisenroth, M. P · 2018
Later among the works it cites.
A general reinforcement learning algorithm that masters chess, shogi, and go through self-play
Silver, D., Hubert, T., Schrittwieser, J., Antonoglou, I., Lai, M., Guez, A., Lanctot, M., Sifre, L., Kumaran, D., Graepel, T., et al · 2018
Later among the works it cites.
Meta-gradient reinforcement learning
Xu, Z., van Hasselt, H., and Silver, D · 2018
Later among the works it cites.
Direct policy transfer via hidden parameter markov decision processes
Yao, J., Killian, T. W., Konidaris, G., and Doshi-Velez, F · 2018
Later among the works it cites.
Fast context adaptation via meta-learning
Zintgraf, L. M., Shiarlis, K., Kurin, V., Hofmann, K., and Whiteson, S · 2018
Later among the works it cites.
Dota 2 with large scale deep reinforcement learning
Berner, C., Brockman, G., Chan, B., Cheung, V., Debiak, P., Dennison, C., Farhi, D., Fischer, Q., Hashme, S., Hesse, C., Józefowicz, R., Gray, S., Olsson, C., Pachocki, J. W., Petrov, M., de Oliveira Pinto, H. P., Raiman, J., Salimans, T., Schlatter, J., Schneider, J., Sidor, S., Sutskever, I., Tang, J., Wolski, F., and Zhang, S · 2019
Later among the works it cites.
Quantifying generalization in reinforcement learning
Cobbe, K., Klimov, O., Hesse, C., Kim, T., and Schulman, J · 2019
Later among the works it cites.
Fast task inference with variational intrinsic successor features
Hansen, S., Dabney, W., Barreto, A., Van de Wiele, T., Warde-Farley, D., and Mnih, V · 2019
Later among the works it cites.
Multi-task deep reinforcement learning with popart
Hessel, M., Soyer, H., Espeholt, L., Czarnecki, W., Schmitt, S., and van Hasselt, H · 2019
Later among the works it cites.
Meta reinforcement learning as task inference
Humplik, J., Galashov, A., Hasenclever, L., Ortega, P. A., Teh, Y. W., and Heess, N · 2019
Later among the works it cites.
Human-level performance in 3d multiplayer games with population-based reinforcement learning
Jaderberg, M., Czarnecki, W. M., Dunning, I., Marris, L., Lever, G., Castaneda, A. G., Beattie, C., Rabinowitz, N. C., Morcos, A. S., Ruderman, A., et al · 2019
Later among the works it cites.
State2vec: Off-policy successor features approximators
Madjiheurem, S. and Toni, L · 2019
Later among the works it cites.
Disentangled skill embeddings for reinforcement learning
Petangoda, J. C., Pascual-Diaz, S., Adam, V., Vrancx, P., and Grau-Moya, J · 2019
Later among the works it cites.
Efficient off-policy meta-reinforcement learning via probabilistic context variables
Rakelly, K., Zhou, A., Finn, C., Levine, S., and Quillen, D · 2019
Later among the works it cites.
Siriwardhana, S., Weerasakera, R., Matthies, D. J., and Nanayakkara, S · 2019
Later among the works it cites.
Grandmaster level in starcraft ii using multi-agent reinforcement learning
Vinyals, O., Babuschkin, I., Czarnecki, W. M., Mathieu, M., Dudzik, A., Chung, J., Choi, D. H., Powell, R., Ewalds, T., Georgiev, P., et al · 2019
Later among the works it cites.
Single episode policy transfer in reinforcement learning
Yang, J., Petersen, B., Zha, H., and Faissol, D · 2019
Later among the works it cites.
Ride: Rewarding impact-driven exploration for procedurally-generated environments
Raileanu, R. and Rocktäschel, T · 2020
Closest in time.
Observational overfitting in reinforcement learning
Song, X., Jiang, Y., Tu, S., Du, Y., and Neyshabur, B · 2020
Closest in time.