Fetching the paper…
Reading the bibliography…
While many recent advances in deep reinforcement learning (RL) rely on model-free methods, model-based approaches remain an alluring prospect for their potential to exploit unsupervised data to learn environment model.
Generalized polynomial approximations in markovian decision processes
Schweitzer, P. J. and Seidmann, A. (1985) · 1985
Earlier work this paper cites.
Integrated architectures for learning, planning, and reacting based on approximating dynamic programming
Sutton, R. S. (1990) · 1990
Earlier work this paper cites.
Issues in using function approximation for reinforcement learning
Thrun, S. and Schwartz, A. (1993) · 1993
Earlier work this paper cites.
A natural policy gradient
Kakade, S. M. (2002) · 2002
Earlier work this paper cites.
A sparse sampling algorithm for near-optimal planning in large markov decision processes
Kearns, M., Mansour, Y., and Ng, A. Y. (2002) · 2002
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
Kearns, M. and Singh, S. (2002) · 2002
Earlier work this paper cites.
Using confidence bounds for exploitation-exploration trade-offs
Auer, P. (2003) · 2003
Earlier work this paper cites.
R-max-a general polynomial time algorithm for near-optimal reinforcement learning
Brafman, R. I. and Tennenholtz, M. (2003) · 2003
Earlier work this paper cites.
Least-squares policy iteration
Lagoudakis, M. G. and Parr, R. (2003) · 2003
Earlier work this paper cites.
Bandit based monte-carlo planning
Kocsis, L. and Szepesvári, C. (2006) · 2006
Earlier work this paper cites.
Learning near-optimal policies with bellman-residual minimization based fitted policy iteration and a single sample path
Antos, A., Szepesvári, C., and Munos, R. (2008) · 2008
Earlier work this paper cites.
A bayesian sampling approach to exploration in reinforcement learning
Asmuth, J., Li, L., Littman, M. L., Nouri, A., and Wingate, D. (2009) · 2009
Earlier work this paper cites.
REGAL: A regularization based algorithm for reinforcement learning in weakly communicating MDPs
Bartlett, P. L. and Tewari, A. (2009) · 2009
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Jaksch, T., Ortner, R., and Auer, P. (2010) · 2010
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Abbasi-Yadkori, Y., Pál, D., and Szepesvári, C. (2011) · 2011
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M. (2013) · 2013
Earlier work this paper cites.
Partial monitoring—classification, regret bounds, and algorithms
Bartók, G., Foster, D. P., Pál, D., Rakhlin, A., and Szepesvári, C. (2014) · 2014
Cited alongside, same era.
Generative adversarial nets
Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y. (2014) · 2014
Cited alongside, same era.
Deep learning for real-time atari game play using offline monte-carlo tree search planning
Guo, X., Singh, S., Lee, H., Lewis, R. L., and Wang, X. (2014) · 2014
Cited alongside, same era.
Conditional generative adversarial nets
Mirza, M. and Osindero, S. (2014) · 2014
Cited alongside, same era.
Deep multi-scale video prediction beyond mean square error
Mathieu, M., Couprie, C., and LeCun, Y. (2015) · 2015
Cited alongside, same era.
Mastering the game of go with deep neural networks and tree search
Silver et al., D. (2016) · 2016
Later among the works it cites.
Deep reinforcement learning with double q-learning
Van Hasselt, H., Guez, A., and Silver, D. (2016) · 2016
Later among the works it cites.
Arjovsky, M., Chintala, S., and Bottou, L. (2017) · 2017
Later among the works it cites.
Noisy networks for exploration
Fortunato, M., Azar, M. G., Piot, B., Menick, J., Osband, I., Graves, A., Mnih, V., Munos, R., Hassabis, D., Pietquin, O., et al. (2017) · 2017
Later among the works it cites.
Improved training of wasserstein gans
Gulrajani, I., Ahmed, F., Arjovsky, M., Dumoulin, V., and Courville, A. C. (2017) · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al. (2015) · 2015
Cited alongside, same era.
Action-conditional video prediction using deep networks in atari games
Oh, J., Guo, X., Lee, H., Lewis, R. L., and Singh, S. (2015) · 2015
Cited alongside, same era.
Trust region policy optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M., and Moritz, P. (2015) · 2015
Cited alongside, same era.
From pixels to torques: Policy learning with deep dynamical models
Wahlström, N., Schön, T. B., and Deisenroth, M. P. (2015) · 2015
Cited alongside, same era.
Exploratory gradient boosting for reinforcement learning in complex domains
Abel, D., Agarwal, A., Diaz, F., Krishnamurthy, A., and Schapire, R. E. (2016) · 2016
Cited alongside, same era.
Unifying count-based exploration and intrinsic motivation
Bellemare, M., Srinivasan, S., Ostrovski, G., Schaul, T., Saxton, D., and Munos, R. (2016) · 2016
Cited alongside, same era.
Openai gym
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W. (2016) · 2016
Cited alongside, same era.
Isola, P., Zhu, J.-Y., Zhou, T., and Efros, A. A. (2017) · 2017
Later among the works it cites.
Machado, M. C., Bellemare, M. G., Talvitie, E., Veness, J., Hausknecht, M., and Bowling, M. (2017) · 2017
Later among the works it cites.
Value prediction network
Oh, J., Singh, S., and Lee, H. (2017) · 2017
Later among the works it cites.
Count-based exploration with neural density models
Ostrovski, G., Bellemare, M. G., Oord, A. v. d., and Munos, R. (2017) · 2017
Later among the works it cites.
Imagination-augmented agents for deep reinforcement learning
Weber, T., Racanière, S., Reichert, D. P., Buesing, L., Guez, A., Rezende, D. J., Badia, A. P., Vinyals, O., Heess, N., Li, Y., et al. (2017) · 2017
Later among the works it cites.
Efficient exploration through bayesian deep q-networks
Azizzadenesheli, K., and Anandkumar, A. (2018) · 2018
Closest in time.
David Ha, J. S. (2018) · 2018
Closest in time.
The effect of planning shape on dyna-style planning in high-dimensional state spaces
Holland, G. Z., Talvitie, E. J., and Bowling, M. (2018) · 2018
Closest in time.
Efficient exploration for dialogue policy learning with bbq networks & replay buffer spiking
Lipton, Z. C., Gao, J., Li, L., Li, X., Ahmed, F., and Deng, L. (2018) · 2018
Closest in time.
Spectral normalization for generative adversarial networks
Miyato, T., Kataoka, T., Koyama, M., and Yoshida, Y. (2018) · 2018
Closest in time.