Fetching the paper…
Reading the bibliography…
Despite their great success, there is still no comprehensive theoretical understanding of learning with Deep Neural Networks (DNNs) or their inner organization.
Stochastic relaxation, gibbs distributions, and the bayesian restoration of images, neurocomputing: foundations of research, 1988
Stuart Geman and Donald Geman · 1988
Earlier work this paper cites.
The Fokker-Planck Equation: Methods of Solution and Applications
H. Risken · 1989
Earlier work this paper cites.
Neural Networks: A Comprehensive Foundation
Simon Haykin · 1998
Earlier work this paper cites.
The information bottleneck method
Naftali Tishby, Fernando C. Pereira, and William Bialek · 1999
Earlier work this paper cites.
Rotation invariant spherical harmonic representation of 3d shape descriptors
Michael Kazhdan, Thomas Funkhouser, and Szymon Rusinkiewicz · 2003
Earlier work this paper cites.
Estimation of entropy and mutual information
Liam Paninski · 2003
Earlier work this paper cites.
Estimating mutual information
Alexander Kraskov, Harald Stögbauer, and Peter Grassberger · 2004
Earlier work this paper cites.
Elements of Information Theory (Wiley Series in Telecommunications and Signal Processing)
Thomas M. Cover and Joy A Thomas · 2006
Earlier work this paper cites.
Exploring strategies for training deep neural networks
Hugo Larochelle, Yoshua Bengio, Jérôme Louradour, and Pascal Lamblin · 2009
Cited alongside, same era.
Improving neural networks by preventing co-adaptation of feature detectors
Geoffrey E Hinton, Nitish Srivastava, Alex Krizhevsky, Ilya Sutskever, and Ruslan R Salakhutdinov · 2012
Cited alongside, same era.
Speech recognition with deep recurrent neural networks
Alex Graves, Abdel-rahman Mohamed, and Geoffrey Hinton · 2013
Cited alongside, same era.
Junghwan Cho, Kyewook Lee, Ellie Shin, Garry Choy, and Synho Do · 2015
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Text understanding from scratch
Xiang Zhang and Yann LeCun · 2015
Later among the works it cites.
Information Dropout: Learning Optimal Representations Through Noisy Computation
A. Achille and S. Soatto · 2016
Later among the works it cites.
Understanding intermediate layers using linear classifier probes, 2016
Guillaume Alain and Yoshua Bengio · 2016
Later among the works it cites.
Optimal architectures in a solvable model of deep networks
Jonathan Kadmon and Haim Sompolinsky · 2016
Later among the works it cites.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Deep learning
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton · 2015
Cited alongside, same era.
Deep learning and the information bottleneck principle
Naftali Tishby and Noga Zaslavsky · 2015
Cited alongside, same era.
David Balduzzi, Marcus Frean, Lennox Leary, J. P. Lewis, Kurt Wan-Duo Ma, and Brian McWilliams · 2017
Closest in time.
Mixing complexity and its applications to neural networks
Michal Moshkovich and Naftali Tishby · 2017
Closest in time.