Fetching the paper…
Reading the bibliography…
Deep neural networks are surprisingly efficient at solving practical tasks, but the theory behind this phenomenon is only starting to catch up with the practice.
Analysis of individual differences in multidimensional scaling via an N-way generalization of “Eckart-Young” decomposition
J Douglas Carroll and Jih-Jie Chang · 1970
Earlier work this paper cites.
Foundations of the parafac procedure: models and conditions for an ”explanatory” multimodal factor analysis
Richard A Harshman · 1970
Earlier work this paper cites.
Approximation by superpositions of a sigmoidal function
George Cybenko · 1989
Earlier work this paper cites.
Multilayer feedforward networks are universal approximators
Kurt Hornik, Maxwell Stinchcombe, and Halbert White · 1989
Earlier work this paper cites.
Handwritten digit recognition with a back-propagation network
Yann LeCun, Bernhard E Boser, John S Denker, Donnie Henderson, Richard E Howard, Wayne E Hubbard, and Lawrence D Jackel · 1990
Earlier work this paper cites.
Basic algebraic geometry , volume 2
Igor Rostislavovich Shafarevich and Kurt Augustus Hirsch · 1994
Earlier work this paper cites.
Convolutional networks for images, speech, and time series
Yann LeCun, Yoshua Bengio, et al · 1995
Earlier work this paper cites.
Learning to forget: Continual prediction with LSTM
Felix A Gers, Jürgen Schmidhuber, and Fred Cummins · 1999
Earlier work this paper cites.
Lectures on analytic differential equations , volume 86
Yu S Ilyashenko and Sergei Yakovenko · 2008
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky and Geoffrey Hinton · 2009
Earlier work this paper cites.
Hierarchical singular value decomposition of tensors
Lars Grasedyck · 2010
Earlier work this paper cites.
Shallow vs. deep sum-product networks
Olivier Delalleau and Yoshua Bengio · 2011
Earlier work this paper cites.
An introduction to hierarchical (H-) rank and TT-rank of tensors with examples
Lars Grasedyck and Wolfgang Hackbusch · 2011
Cited alongside, same era.
Extensions of recurrent neural network language model
Tomáš Mikolov, Stefan Kombrink, Lukáš Burget, Jan Černockỳ, and Sanjeev Khudanpur · 2011
Cited alongside, same era.
Tensor-train decomposition
Ivan V Oseledets · 2011
Cited alongside, same era.
Scikit-learn: Machine learning in Python
F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay · 2011
Cited alongside, same era.
Tensor spaces and numerical tensor calculus , volume 42
Wolfgang Hackbusch · 2012
Cited alongside, same era.
Speech recognition with deep recurrent neural networks
Alex Graves, Abdel-rahman Mohamed, and Geoffrey Hinton · 2013
Going deeper with convolutions
Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich · 2015
Later among the works it cites.
Convolutional rectifier networks as generalized tensor decompositions
Nadav Cohen and Amnon Shashua · 2016
Later among the works it cites.
On the expressive power of deep learning: A tensor analysis
Nadav Cohen, Or Sharir, and Amnon Shashua · 2016
Later among the works it cites.
The power of depth for feedforward neural networks
Ronen Eldan and Ohad Shamir · 2016
Later among the works it cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Later among the works it cites.
Alexander Novikov, Mikhail Trofimov, and Ivan Oseledets · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
On the importance of initialization and momentum in deep learning
Ilya Sutskever, James Martens, George Dahl, and Geoffrey Hinton · 2013
Cited alongside, same era.
On the expressive efficiency of sum product networks
James Martens and Venkatesh Medabalimi · 2014
Cited alongside, same era.
On the number of linear regions of deep neural networks
Guido F Montufar, Razvan Pascanu, Kyunghyun Cho, and Yoshua Bengio · 2014
Cited alongside, same era.
TensorFlow: Large-scale machine learning on heterogeneous systems, 2015
Martín Abadi, Ashish Agarwal, Paul Barham, Eugene Brevdo, Zhifeng Chen, Craig Citro, Greg S. Corrado, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Ian Goodfellow, Andrew Harp, Geoffrey Irving, Michael Isard, Yangqing Jia, Rafal Jozefowicz, Lukasz Kaiser, Manjunath Kudlur, Josh Levenberg, Dan Mané, Rajat Monga, Sherry Moore, Derek Murray, Chris Olah, Mike Schuster, Jonathon Shlens, Benoit Steiner, Ilya Sutskever, Kunal Talwar, Paul Tucker, Vincent Vanhoucke, Vijay Vasudevan, Fernanda Viégas, Oriol Vinyals, Pete Warden, Martin Wattenberg, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng · 2015
Cited alongside, same era.
The Hackbusch conjecture on tensor formats
Weronika Buczyńska, Jarosław Buczyński, and Mateusz Michałek · 2015
Cited alongside, same era.
Later among the works it cites.
Supervised learning with tensor networks
Edwin Stoudenmire and David J Schwab · 2016
Later among the works it cites.
On multiplicative integration with recurrent neural networks
Yuhuai Wu, Saizheng Zhang, Ying Zhang, Yoshua Bengio, and Ruslan R Salakhutdinov · 2016
Later among the works it cites.
On the expressive power of deep neural networks
Maithra Raghu, Ben Poole, Jon Kleinberg, Surya Ganguli, and Jascha Sohl-Dickstein · 2017
Closest in time.
Opening the black box of deep neural networks via information
Ravid Shwartz-Ziv and Naftali Tishby · 2017
Closest in time.
Long-term forecasting using tensor-train RNNs
Rose Yu, Stephan Zheng, Anima Anandkumar, and Yisong Yue · 2017
Closest in time.