Finite-sample analysis of bellman residual minimization
Maillard, O.-A., Munos, R., Lazaric, A., and Ghavamzadeh, M · 2010
Cited alongside, same era.
Should one compute the temporal difference fix point or minimize the bellman residual? the unified oblique projection view
Scherrer, B · 2010
Cited alongside, same era.
Understanding machine learning: From theory to algorithms
Shalev-Shwartz, S. and Ben-David, S · 2014
Cited alongside, same era.
Continuous control with deep reinforcement learning
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., Petersen, S., Beattie, C., Sadik, A., Antonoglou, I., King, H., Kumaran, D., Wierstra, D., Legg, S., and Hassabis, D · 2015
Cited alongside, same era.
Prioritized experience replay
Schaul, T., Quan, J., Antonoglou, I., and Silver, D · 2015
Cited alongside, same era.
On-policy vs. off-policy updates for deep reinforcement learning
Hausknecht, M. and Stone, P · 2016
Cited alongside, same era.
Safe and efficient off-policy reinforcement learning
Munos, R., Stepleton, T., Harutyunyan, A., and Bellemare, M · 2016
Cited alongside, same era.
An emphatic approach to the problem of off-policy temporal-difference learning
Sutton, R. S., Mahmood, A. R., and White, M · 2016
Cited alongside, same era.
Wasserstein generative adversarial networks
Arjovsky, M., Chintala, S., and Bottou, L · 2017
Cited alongside, same era.
Is the bellman residual a bad proxy?
Geist, M., Piot, B., and Pietquin, O · 2017
Cited alongside, same era.
Reinforcement learning with deep energy-based policies
Haarnoja, T., Tang, H., Abbeel, P., and Levine, S · 2017
Cited alongside, same era.