Shallow vs. deep sum-product networks
O. Delalleau and Y. Bengio · 2011
Later among the works it cites.
Representation learning: A review and new perspectives
Y. Bengio, A. Courville, and P. Vincent · 2013
Later among the works it cites.
Intriguing properties of neural networks
C. Szegedy, W. Zaremba, I. Sutskever, J. Bruna, D. Erhan, I. Goodfellow, and R. Fergus · 2014
Later among the works it cites.
Language models for image captioning: The quirks and what works
J. Devlin, H. Cheng, H. Fang, S. Gupta, L. Deng, X. He, G. Zweig, and M. Mitchell · 2015
Later among the works it cites.
On the expressive power of deep learning: A tensor analysis
N. Cohen, O. Sharir, and A. Shashua · 2016
Later among the works it cites.
Deep Learning
I. Goodfellow, Y. Bengio, and A. Courville · 2016
Later among the works it cites.
“Why should I trust you?”: Explaining the predictions of any classifier
M. T. Ribeiro, S. Singh, and C. Guestrin · 2016
Later among the works it cites.
Integration methods and optimization algorithms
D. Scieur, V. Roulet, F. Bach, and A. d’Aspremont · 2017
Later among the works it cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, N. Aidan, L. Kaiser, and I. Polosukhin · 2017
Later among the works it cites.
Understanding deep learning requires rethinking generalization
C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals · 2017
Later among the works it cites.
Neural tangent kernel: Convergence and generalization in neural networks
A. Jacot, F. Gabriel, and C. Hongler · 2018
Later among the works it cites.
Visual interpretability for deep learning: A survey
Q. Zhang and S.-C. Zhu · 2018
Later among the works it cites.
Approximating CNNs with bag-of-local-features models works surprisingly well on ImageNet
W. Brendel and M. Bethge · 2019
Later among the works it cites.