Fetching the paper…
Reading the bibliography…
One of the key approaches to save samples in reinforcement learning (RL) is to use knowledge from an approximate model such as its simulator.
Sample complexity of reinforcement learning using linearly combined model ensembles
Modi, A., Jiang, N., Tewari, A., and Singh, S. (2019) · 1910
Earlier work this paper cites.
Reinforcement learning is direct adaptive optimal control
Sutton, R. S., Barto, A. G., and Williams, R. J. (1992) · 1992
Earlier work this paper cites.
Q-learning
Watkins, C. J. and Dayan, P. (1992) · 1992
Earlier work this paper cites.
Soccer server: A simulator for RoboCup
Itsuki, N. (1995) · 1995
Earlier work this paper cites.
Sailing strategies, an application involving stochastics, optimization, and statistics
Vanderbei, R. (1996) · 1996
Earlier work this paper cites.
An approach to lifelong reinforcement learning through multiple environments
Tanaka, F. and Yamamura, M. (1997) · 1997
Earlier work this paper cites.
Multi-agent reinforcement learning for traffic light control
Wiering, M. (2000) · 2000
Earlier work this paper cites.
Online learning control by association and reinforcement
Si, J. and Wang, Y.-T. (2001) · 2001
Earlier work this paper cites.
Meta-learning in reinforcement learning
Schweighofer, N. and Doya, K. (2003) · 2003
Earlier work this paper cites.
The sample complexity of exploration in the multi-armed bandit problem
Mannor, S. and Tsitsiklis, J. N. (2004) · 2004
Earlier work this paper cites.
Autonomous inverted helicopter flight via reinforcement learning
Ng, A. Y., Coates, A., Diel, M., Ganapathi, V., Schulte, J., Tse, B., Berger, E., and Liang, E. (2006) · 2006
Earlier work this paper cites.
Multi-task reinforcement learning: A hierarchical Bayesian approach
Wilson, A., Fern, A., Ray, S., and Tadepalli, P. (2007) · 2007
Earlier work this paper cites.
Transfer learning for reinforcement learning domains: A survey
Taylor, M. E. and Stone, P. (2009) · 2009
Cited alongside, same era.
Bayesian Multi-Task Reinforcement Learning
Lazaric, A. and Ghavamzadeh, M. (2010) · 2010
Cited alongside, same era.
Transfer in reinforcement learning: A framework and a survey
Lazaric, A. (2012) · 2012
Cited alongside, same era.
Directed exploration in reinforcement learning with transferred knowledge
Mann, T. A. and Choe, Y. (2012) · 2012
Cited alongside, same era.
Minimax PAC bounds on the sample complexity of reinforcement learning with a generative model
Azar, M. G., Munos, R., and Kappen, H. J. (2013) · 2013
Cited alongside, same era.
Sample complexity of multi-task reinforcement learning
Brunskill, E. and Li, L. (2013) · 2013
Cited alongside, same era.
Scalable lifelong reinforcement learning
Zhan, Y., Ammar, H. B., and Taylor, M. E. (2017) · 2017
Later among the works it cites.
Policy and value transfer in lifelong reinforcement learning
Abel, D., Jinnai, Y., Guo, S. Y., Konidaris, G., and Littman, M. (2018) · 2018
Later among the works it cites.
Continuous adaptation via meta-learning in nonstationary and competitive environments
Al-Shedivat, M., Bansal, T., Burda, Y., Sutskever, I., Mordatch, I., and Abbeel, P. (2018) · 2018
Later among the works it cites.
Meta-reinforcement learning of structured exploration strategies
Gupta, A., Mendonca, R., Liu, Y., Abbeel, P., and Levine, S. (2018) · 2018
Later among the works it cites.
PAC reinforcement learning with an imperfect model
Jiang, N. (2018) · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Online multi-task learning for policy gradient methods
Ammar, H. B., Eaton, E., Ruvolo, P., and Taylor, M. (2014) · 2014
Cited alongside, same era.
PAC-inspired option discovery in lifelong reinforcement learning
Brunskill, E. and Li, L. (2014) · 2014
Cited alongside, same era.
Sparse multi-task reinforcement learning
Calandriello, D., Lazaric, A., and Restelli, M. (2014) · 2014
Cited alongside, same era.
Mastering the game of go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al. (2016) · 2016
Cited alongside, same era.
Hindsight experience replay
Andrychowicz, M., Wolski, F., Ray, A., Schneider, J., Fong, R., Welinder, P., McGrew, B., Tobin, J., Pieter Abbeel, O., and Zaremba, W. (2017) · 2017
Cited alongside, same era.
CARLA: An open urban driving simulator
Dosovitskiy, A., Ros, G., Codevilla, F., Lopez, A., and Koltun, V. (2017) · 2017
Cited alongside, same era.
Sæmundsson, S., Hofmann, K., and Deisenroth, M. P. (2018) · 2018
Later among the works it cites.
Near-optimal time and sample complexities for solving Markov decision processes with a generative model
Sidford, A., Wang, M., Wu, X., Yang, L., and Ye, Y. (2018) · 2018
Later among the works it cites.
Probability: Theory and examples
Durrett, R. (2019) · 2019
Closest in time.
A guide to deep learning in healthcare
Esteva, A., Robicquet, A., Ramsundar, B., Kuleshov, V., DePristo, M., Chou, K., Cui, C., Corrado, G., Thrun, S., and Dean, J. (2019) · 2019
Closest in time.
Efficient off-policy meta-reinforcement learning via probabilistic context variables
Rakelly, K., Zhou, A., Finn, C., Levine, S., and Quillen, D. (2019) · 2019
Closest in time.
Model-based reinforcement learning with value-targeted regression
Ayoub, A., Jia, Z., Szepesvari, C., Wang, M., and Yang, L. F. (2020) · 2020
Closest in time.
Transfer learning
Yang, Q., Zhang, Y., Dai, W., and Pan, S. J. (2020) · 2020
Closest in time.