Robust stochastic approximation approach to stochastic programming
Nemirovski, A., Juditsky, A., Lan, G., and Shapiro, A · 2009
Cited alongside, same era.
Rectified linear units improve restricted boltzmann machines
Nair, V. and Hinton, G. E · 2010
Cited alongside, same era.
Gradient temporal-difference learning algorithms
Maei, H. R · 2011
Cited alongside, same era.
Horde: A scalable real-time architecture for learning knowledge from unsupervised sensorimotor interaction
Sutton, R. S., Modayil, J., Delp, M., Degris, T., Pilarski, P. M., White, A., and Precup, D · 2011
Cited alongside, same era.
Matrix analysis (2nd Edition)
Horn, R. A. and Johnson, C. R · 2012
Cited alongside, same era.
Markov decision processes: discrete stochastic dynamic programming
Puterman, M. L · 2014
Cited alongside, same era.
Lectures on stochastic programming: modeling and theory
Shapiro, A., Dentcheva, D., and Ruszczyński, A · 2014
Cited alongside, same era.
Doubly robust off-policy value evaluation for reinforcement learning
Original
Jiang, N. and Li, L · 2015
Cited alongside, same era.
Finite-sample analysis of proximal gradient td algorithms
Liu, B., Liu, J., Ghavamzadeh, M., Mahadevan, S., and Petrik, M · 2015
Cited alongside, same era.
Openai gym
Original
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W · 2016
Cited alongside, same era.
Learning from conditional distributions via dual embeddings
Original
Dai, B., He, N., Pan, Y., Boots, B., and Song, L · 2016
Cited alongside, same era.
Fast gradient-descent methods for temporal-difference learning with linear function approximation
Sutton, R. S., Maei, H. R., Precup, D., Bhatnagar, S., Silver, D., Szepesvári, C., and Wiewiora, E
Cited in the paper.