R. S. Sutton, “Learning to predict by the methods of temporal differences,” Machine Learning , vol. 3, no. 1, pp. 9–44, 1988
1988
Earlier work this paper cites.
G. Tesauro, “Temporal difference learning and td-gammon,” Communications of the ACM , vol. 38, no. 3, pp. 58–68, 1995
1995
Earlier work this paper cites.
R. S. Sutton and A. G. Barto, Introduction to Reinforcement Learning , 1st ed. MIT Press, 1998
1998
Earlier work this paper cites.
F. Elizalde, L. Sucar, M. Luque, F. Díez, and A. Reyes Ballesteros, “Policy explanation in factored markov decision processes,” in European Workshop on Probabilistic Graphical Models , 2008
2008
Earlier work this paper cites.
L. v. d. Maaten and G. Hinton, “Visualizing data using t-sne,” Journal of Machine Learning Research , vol. 9, pp. 2579–2605, 2008
2008
Earlier work this paper cites.
O. Z. Khan, P. Poupart, and J. P. Black, “Minimal sufficient explanations for factored markov decision processes,” in International Conference on Automated Planning and Scheduling , 2009
2009
Earlier work this paper cites.
V. Nair and G. E. Hinton, “Rectified linear units improve restricted boltzmann machines,” in Proceedings of the 27th International Conference on Machine Learning (ICML) , 2010
2010
Earlier work this paper cites.
T. Dodson, N. Mattei, and J. Goldsmith, “A natural language argumentation interface for explanation generation in markov decision processes,” in Algorithmic Decision Theory , 2011
2011
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” in Advances in Neural Information Processing Systems (NeurIPS) , 2012
2012
Earlier work this paper cites.
E. Todorov, T. Erez, and Y. Tassa, “Mujoco: A physics engine for model-based control,” IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) , pp. 5026–5033, 2012
2012
Earlier work this paper cites.
V. Mnih, K. Kavukcuoglu, D. Silver, A. Graves, I. Antonoglou, D. Wierstra, and M. Riedmiller, “Playing atari with deep reinforcement learning,” in Advances in Neural Information Processing Systems (NeurIPS) Deep Learning Workshop , 2013
2013
Earlier work this paper cites.
M. G. Bellemare, Y. Naddaf, J. Veness, and M. Bowling, “The arcade learning environment: An evaluation platform for general agents,” Journal of Artificial Intelligence Research , vol. 47, pp. 253–279, 2013
2013
Earlier work this paper cites.
K. Simonyan, A. Vedaldi, and A. Zisserman, “Deep inside convolutional networks: Visualising image classification models and saliency maps,” in International Conference on Learning Representations (ICLR) , 2014
2014
Earlier work this paper cites.
M. D. Zeiler and R. Fergus, “Visualizing and understanding convolutional networks,” in The European Conference on Computer Vision (ECCV) , 2014
2014
Earlier work this paper cites.
K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale image recognition,” arXiv preprint arXiv:1409.1556 , 2014
Original
2014
Earlier work this paper cites.
J. T. Springenberg, A. Dosovitskiy, T. Brox, and M. Riedmiller, “Striving for simplicity: The all convolutional net,” in International Conference on Learning Representations (ICLR) , 2015
2015
Earlier work this paper cites.
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, S. Petersen, C. Beattie, A. Sadik, I. Antonoglou, H. King, D. Kumaran, D. Wierstra, S. Legg, and D. Hassabis, “Human-level control through deep reinforcement learning,” Nature , vol. 518, no. 7540, pp. 529–533, 2015
2015
Earlier work this paper cites.
I. Sorokin, A. Seleznev, M. Pavlov, A. Fedorov, and A. Ignateva, “Deep attention recurrent q-network,” in Deep Reinforcement Learning Workshop, Advances in Neural Information Processing Systems (NeurIPS) , 2015
2015
Earlier work this paper cites.
K. Xu, J. Ba, R. Kiros, K. Cho, A. Courville, R. Salakhutdinov, R. Zemel, and Y. Bengio, “Show, attend and tell: Neural image caption generation with visual attention,” in Proceedings of the 32nd International Conference on Machine Learning (ICML) , 2015
2015
Earlier work this paper cites.