Acting optimally in partially observable stochastic domains
Cassandra, A. R., Kaelbling, L. P., and Littman, M. L · 1994
Earlier work this paper cites.
Long short-term memory
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
Predictive representations of state
Littman, M. L., Sutton, R. S., and Singh, S. P · 2001
Earlier work this paper cites.
Learning predictive state representations
Singh, S. P., Littman, M. L., Jong, N. K., Pardoe, D., and Stone, P · 2003
Earlier work this paper cites.
Rectified linear units improve restricted boltzmann machines
Nair, V. and Hinton, G. E · 2010
Earlier work this paper cites.
Algorithms for reinforcement learning
Szepesvári, C · 2010
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M · 2013
Earlier work this paper cites.
Sample complexity of multi-task reinforcement learning
Brunskill, E. and Li, L · 2013
Earlier work this paper cites.
Auto-encoding variational bayes
Kingma, D. P. and Welling, M · 2014
Earlier work this paper cites.
TensorFlow: Large-scale machine learning on heterogeneous systems, 2015
Abadi, M., Agarwal, A., Barham, P., Brevdo, E., Chen, Z., Citro, C., Corrado, G. S., Davis, A., Dean, J., Devin, M., Ghemawat, S., Goodfellow, I., Harp, A., Irving, G., Isard, M., Jia, Y., Jozefowicz, R., Kaiser, L., Kudlur, M., Levenberg, J., Mané, D., Monga, R., Moore, S., Murray, D., Olah, C., Schuster, M., Shlens, J., Steiner, B., Sutskever, I., Talwar, K., Tucker, P., Vanhoucke, V., Vasudevan, V., Viégas, F., Vinyals, O., Warden, P., Wattenberg, M., Wicke, M., Yu, Y., and Zheng, X · 2015
Earlier work this paper cites.
DRAW: A recurrent neural network for image generation
Gregor, K., Danihelka, I., Graves, A., Rezende, D. J., and Wierstra, D · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Earlier work this paper cites.
Action-conditional video prediction using deep networks in Atari games
Oh, J., Guo, X., Lee, H., Lewis, R. L., and Singh, S. P · 2015
Earlier work this paper cites.
DeepMind lab, 2016
Beattie, C., Leibo, J. Z., Teplyashin, D., Ward, T., Wainwright, M., Küttler, H., Lefrancq, A., Green, S., Valdés, V., Sadik, A., Schrittwieser, J., Anderson, K., York, S., Cant, M., Cain, A., Bolton, A., Gaffney, S., King, H., Hassabis, D., Legg, S., and Petersen, S · 2016
Earlier work this paper cites.