Fetching the paper…
Reading the bibliography…
Reinforcement learning (RL) methods have been shown to be capable of learning intelligent behavior in rich domains.
Contextual markov decision processes using generalized linear models
Modi, A. and Tewari, A. (2019) · 1903
Earlier work this paper cites.
Reinforcement leaning in feature space: Matrix bandit, kernels, and regret bound
Yang, L. F. and Wang, M. (2019b) · 1905
Earlier work this paper cites.
Model selection for contextual bandits
Foster, D. J., Krishnamurthy, A., and Luo, H. (2019) · 1906
Earlier work this paper cites.
Provably efficient reinforcement learning with linear function approximation
Jin, C., Yang, Z., Wang, Z., and Jordan, M. I. (2019) · 1907
Earlier work this paper cites.
Approximations of dynamic programs, i
Whitt, W. (1978) · 1978
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
Auer, P., Cesa-Bianchi, N., and Fischer, P. (2002) · 2002
Earlier work this paper cites.
Multiple model-based reinforcement learning
Doya, K., Samejima, K., Katagiri, K.-i., and Kawato, M. (2002) · 2002
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
Kearns, M. and Singh, S. (2002) · 2002
Earlier work this paper cites.
Equivalence notions and model minimization in markov decision processes
Givan, R., Dean, T., and Greig, M. (2003) · 2003
Earlier work this paper cites.
Towards a unified theory of state abstraction for mdps
Li, L., Walsh, T. J., and Littman, M. L. (2006) · 2006
Earlier work this paper cites.
Matrix regularization techniques for online multitask learning
Agarwal, A., Rakhlin, A., and Bartlett, P. (2008) · 2008
Earlier work this paper cites.
On the sample complexity of reinforcement learning with a generative model
Azar, M. G., Munos, R., and Kappen, H. J. (2012) · 2012
Cited alongside, same era.
Mujoco: A physics engine for model-based control
Todorov, E., Erez, T., and Tassa, Y. (2012) · 2012
Cited alongside, same era.
Online learning in mdps with side information
Abbasi-Yadkori, Y. and Neu, G. (2014) · 2014
Cited alongside, same era.
Abstraction selection in model-based reinforcement learning
Jiang, N., Kulesza, A., and Singh, S. (2015) · 2015
Cited alongside, same era.
Continuous control with deep reinforcement learning
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D. (2015) · 2015
Cited alongside, same era.
Contextual decision processes with low bellman rank are pac-learnable
Jiang, N., Krishnamurthy, A., Agarwal, A., Langford, J., and Schapire, R. E. (2017) · 2017
Later among the works it cites.
Epopt: Learning robust neural network policies using model ensembles
Rajeswaran, A., Ghotra, S., Ravindran, B., and Levine, S. (2017) · 2017
Later among the works it cites.
Domain randomization for transferring deep neural networks from simulation to the real world
Tobin, J., Fong, R., Ray, A., Schneider, J., Zaremba, W., and Abbeel, P. (2017) · 2017
Later among the works it cites.
Learning dexterous in-hand manipulation
Andrychowicz, M., Baker, B., Chociej, M., Jozefowicz, R., McGrew, B., Pachocki, J., Petron, A., Plappert, M., Powell, G., Ray, A., et al. (2018) · 2018
Later among the works it cites.
Sample-efficient reinforcement learning with stochastic ensemble value expansion
Buckman, J., Hafner, D., Tucker, G., Brevdo, E., and Lee, H. (2018) · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al. (2015) · 2015
Cited alongside, same era.
Transfer from simulation to real world through learning deep inverse dynamics model
Christiano, P., Shah, Z., Mordatch, I., Schneider, J., Blackwell, T., Tobin, J., Abbeel, P., and Zaremba, W. (2016) · 2016
Cited alongside, same era.
Pac reinforcement learning with rich observations
Krishnamurthy, A., Agarwal, A., and Langford, J. (2016) · 2016
Cited alongside, same era.
Combining model-based policy search with online model learning for control of physical humanoids
Mordatch, I., Mishra, N., Eppner, C., and Abbeel, P. (2016) · 2016
Cited alongside, same era.
Mastering the game of go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al. (2016) · 2016
Cited alongside, same era.
Sample-optimal parametric q-learning using linearly additive features
Yang, L. and Wang, M. (2019a)
Cited in the paper.
On oracle-efficient pac rl with rich observations
Dann, C., Jiang, N., Krishnamurthy, A., Agarwal, A., Langford, J., and Schapire, R. E. (2018) · 2018
Later among the works it cites.
Model-ensemble trust-region policy optimization
Kurutach, T., Clavera, I., Duan, Y., Tamar, A., and Abbeel, P. (2018) · 2018
Later among the works it cites.
Markov decision processes with continuous side information
Modi, A., Jiang, N., Singh, S., and Tewari, A. (2018) · 2018
Later among the works it cites.
Bayesian policy optimization for model uncertainty
Lee, G., Hou, B., Mandalika, A., Lee, J., and Srinivasa, S. S. (2019) · 2019
Closest in time.
Model-based rl in contextual decision processes: Pac bounds and exponential improvements over model-free approaches
Sun, W., Jiang, N., Krishnamurthy, A., Agarwal, A., and Langford, J. (2019) · 2019
Closest in time.