Eligibility traces for off-policy policy evaluation
Precup, D. (2000) · 2000
Earlier work this paper cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Original
Levine, S., Kumar, A., Tucker, G., and Fu, J. (2020) · 2005
Earlier work this paper cites.
Learning near-optimal policies with bellman-residual minimization based fitted policy iteration and a single sample path
Antos, A., Szepesvári, C., and Munos, R. (2008) · 2008
Earlier work this paper cites.
Introduction to Nonparametric Estimation
Tsybakov, A. B. (2009) · 2009
Earlier work this paper cites.
What are the statistical limits of offline rl with linear function approximation?
Original
Wang, R., Foster, D. P., and Kakade, S. M. (2020) · 2010
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Abbasi-Yadkori, Y., Pál, D., and Szepesvári, C. (2011) · 2011
Earlier work this paper cites.
Contextual bandits with linear payoff functions
Chu, W., Li, L., Reyzin, L., and Schapire, R. (2011) · 2011
Earlier work this paper cites.
Batch reinforcement learning
Lange, S., Gabel, T., and Riedmiller, M. (2012) · 2012
Earlier work this paper cites.
Reinforcement learning in robotics: A survey
Kober, J., Bagnell, J. A., and Peters, J. (2013) · 2013
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Original
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D. (2015) · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al. (2015) · 2015
Earlier work this paper cites.
Trust region policy optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M., and Moritz, P. (2015) · 2015
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al. (2016) · 2016
Earlier work this paper cites.
Minimax regret bounds for reinforcement learning
Azar, M. G., Osband, I., and Munos, R. (2017) · 2017
Earlier work this paper cites.
Contextual decision processes with low Bellman rank are PAC-learnable
Jiang, N., Krishnamurthy, A., Agarwal, A., Langford, J., and Schapire, R. E. (2017) · 2017
Earlier work this paper cites.
Agile autonomous driving using end-to-end deep imitation learning
Original
Pan, Y., Cheng, C.-A., Saigol, K., Lee, K., Yan, X., Theodorou, E., and Boots, B. (2017) · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Original
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O. (2017) · 2017
Earlier work this paper cites.