Fetching the paper…
Reading the bibliography…
Transfer and adaptation to new unknown environmental dynamics is a key challenge for reinforcement learning (RL).
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J. (1992) · 1992
Earlier work this paper cites.
Multitask learning
Caruana, R. (1997) · 1997
Earlier work this paper cites.
Planning and acting in partially observable stochastic domains
Kaelbling, L. P., Littman, M. L., and Cassandra, A. R. (1998) · 1998
Earlier work this paper cites.
Module-based reinforcement learning: Experiments with a real robot
Kalmár, Z., Szepesvári, C., and Lőrincz, A. (1998) · 1998
Earlier work this paper cites.
Reinforcement learning: An introduction
Sutton, R. S., Barto, A. G., et al · 1998
Earlier work this paper cites.
ε \varepsilon -mdps: Learning in varying environments
Szita, I., Takács, B., and Lörincz, A. (2002) · 2002
Earlier work this paper cites.
Dynamic multidrug therapies for hiv: Optimal and sti control approaches
Adams, B. M., Banks, H. T., Kwon, H.-D., and Tran, H. T. (2004) · 2004
Earlier work this paper cites.
Transfer learning for reinforcement learning domains: A survey
Taylor, M. E. and Stone, P. (2009) · 2009
Earlier work this paper cites.
Learning parameterized skills
Da Silva, B. C., Konidaris, G., and Barto, A. G. (2012) · 2012
Earlier work this paper cites.
Pharmacogenomics knowledge for personalized medicine
Whirl-Carrillo, M., McDonagh, E. M., Hebert, J., Gong, L., Sangkuhl, K., Thorn, C., Altman, R. B., and Klein, T. E. (2012) · 2012
Earlier work this paper cites.
Auto-encoding variational bayes
Kingma, D. P. and Welling, M. (2014) · 2014
Earlier work this paper cites.
Hidden parameter markov decision processes: an emerging paradigm for modeling families of related tasks
Konidaris, G. and Doshi-Velez, F. (2014) · 2014
Earlier work this paper cites.
Personalized whole-cell kinetic models of metabolism for discovery in genomics and pharmacodynamics
Bordbar, A., McCloskey, D., Zielinski, D. C., Sonnenschein, N., Jamshidi, N., and Palsson, B. O. (2015) · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Earlier work this paper cites.
Universal value function approximators
Schaul, T., Horgan, D., Gregor, K., and Silver, D. (2015) · 2015
Cited alongside, same era.
Hidden parameter markov decision processes: A semiparametric regression approach for discovering latent task parametrizations
Doshi-Velez, F. and Konidaris, G. (2016) · 2016
Cited alongside, same era.
Precision medicine
Hodson, R. (2016) · 2016
Cited alongside, same era.
Prioritized experience replay
Schaul, T., Quan, J., Antonoglou, I., and Silver, D. (2016) · 2016
Cited alongside, same era.
Deep reinforcement learning with double q-learning
Van Hasselt, H., Guez, A., and Silver, D. (2016) · 2016
Cited alongside, same era.
Model-agnostic meta-learning for fast adaptation of deep networks
Finn, C., Abbeel, P., and Levine, S. (2017) · 2017
Cited alongside, same era.
Efficient transfer learning and online adaptation with latent variable models for continuous control
Perez, C. F., Such, F. P., and Karaletsos, T. (2018) · 2018
Later among the works it cites.
Transfer of value functions via variational methods
Tirinzoni, A., Sanchez, R. R., and Restelli, M. (2018) · 2018
Later among the works it cites.
Towards multi-drug adaptive therapy
West, J., You, L., Brown, J., Newton, P. K., and Anderson, A. R. A. (2018) · 2018
Later among the works it cites.
Learning to explore with meta-policy gradient
Xu, T., Liu, Q., Zhao, L., Xu, W., and Peng, J. (2018) · 2018
Later among the works it cites.
Direct policy transfer via hidden parameter markov decision processes
Yao, J., Killian, T., Konidaris, G., and Doshi-Velez, F. (2018) · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
beta-vae: Learning basic visual concepts with a constrained variational framework
Higgins, I., Matthey, L., Pal, A., Burgess, C., Glorot, X., Botvinick, M., Mohamed, S., and Lerchner, A. (2017) · 2017
Cited alongside, same era.
Robust and efficient transfer learning with hidden parameter markov decision processes
Killian, T. W., Daulton, S., Konidaris, G., and Doshi-Velez, F. (2017) · 2017
Cited alongside, same era.
Epopt: Learning robust neural network policies using model ensembles
Rajeswaran, A., Ghotra, S., Ravindran, B., and Levine, S. (2017) · 2017
Cited alongside, same era.
Preparing for the unknown: Learning a universal policy with online system identification
Yu, W., Tan, J., Liu, C. K., and Turk, G. (2017) · 2017
Cited alongside, same era.
Integrating evolutionary dynamics into treatment of metastatic castrate-resistant prostate cancer
Zhang, J., Cunningham, J. J., Brown, J. S., and Gatenby, R. A. (2017) · 2017
Cited alongside, same era.
Self-consistent trajectory autoencoder: Hierarchical reinforcement learning with trajectory embeddings
Co-Reyes, J. D., Liu, Y., Gupta, A., Eysenbach, B., Abbeel, P., and Levine, S. (2018) · 2018
Cited alongside, same era.
Zhang, A., Satija, H., and Pineau, J. (2018) · 2018
Later among the works it cites.
An overview of machine teaching
Zhu, X., Singla, A., Zilles, S., and Rafferty, A. N. (2018) · 2018
Later among the works it cites.
Vpe: Variational policy embedding for transfer reinforcement learning
Arnekvist, I., Kragic, D., and Stork, J. A. (2019) · 2019
Closest in time.
Fingerprint policy optimisation for robust reinforcement learning
Paul, S., Osborne, M. A., and Whiteson, S. (2019) · 2019
Closest in time.
Deep reinforcement learning and simulation as a path toward precision medicine
Petersen, B. K., Yang, J., Grathwohl, W. S., Cockrell, C., Santiago, C., An, G., and Faissol, D. M. (2019) · 2019
Closest in time.
Efficient off-policy meta-reinforcement learning via probabilistic context variables
Rakelly, K., Zhou, A., Finn, C., Levine, S., and Quillen, D. (2019) · 2019
Closest in time.
Meta-learning with latent embedding optimization
Rusu, A. A., Rao, D., Sygnowski, J., Vinyals, O., Pascanu, R., Osindero, S., and Hadsell, R. (2019) · 2019
Closest in time.
Fast context adaptation via meta-learning
Zintgraf, L. M., Shiarlis, K., Kurin, V., Hofmann, K., and Whiteson, S. (2019) · 2019
Closest in time.