A markovian decision process
Bellman, R · 1957
Earlier work this paper cites.
Error bounds for approximate policy iteration
Munos, R · 2003
Earlier work this paper cites.
Convex optimization
Boyd, S., Boyd, S. P., and Vandenberghe, L · 2004
Earlier work this paper cites.
Error bounds for approximate value iteration
Munos, R · 2005
Earlier work this paper cites.
Conservative q-learning for offline reinforcement learning
Original
Kumar, A., Zhou, A., Tucker, G., and Levine, S · 2006
Earlier work this paper cites.
Accelerating online reinforcement learning with offline datasets
Original
Nair, A., Dalal, M., Gupta, A., and Levine, S · 2006
Earlier work this paper cites.
Curriculum learning
Bengio, Y., Louradour, J., Collobert, R., and Weston, J · 2009
Earlier work this paper cites.
Error propagation for approximate policy and value iteration
Farahmand, A. M., Munos, R., and Szepesvári, C · 2010
Earlier work this paper cites.
Deep auto-encoder neural networks in reinforcement learning
Lange, S. and Riedmiller, M · 2010
Earlier work this paper cites.
Counterfactual reasoning and learning systems: The example of computational advertising
Bottou, L., Peters, J., Quiñonero-Candela, J., Charles, D. X., Chickering, D. M., Portugaly, E., Ray, D., Simard, P., and Snelson, E · 2013
Earlier work this paper cites.
Guided policy search
Levine, S. and Koltun, V · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Original
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Scheduled sampling for sequence prediction with recurrent neural networks
Original
Bengio, S., Vinyals, O., Jaitly, N., and Shazeer, N · 2015
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Original
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., and Kavukcuoglu, K · 2016
Earlier work this paper cites.
Deep reinforcement learning with double q-learning
Van Hasselt, H., Guez, A., and Silver, D · 2016
Earlier work this paper cites.
Dueling network architectures for deep reinforcement learning
Wang, Z., Schaul, T., Hessel, M., Hasselt, H., Lanctot, M., and Freitas, N · 2016
Earlier work this paper cites.
Hindsight Experience Replay
Andrychowicz, M., Wolski, F., Ray, A., Schneider, J., Fong, R., Welinder, P., McGrew, B., Tobin, J., Pieter Abbeel, O., and Zaremba, W · 2017
Earlier work this paper cites.