Fetching the paper…
Reading the bibliography…
Using established principles from Statistics and Information Theory, we show that invariance to nuisance factors in a deep neural network is equivalent to information minimality of the learned representation, and that stacking layers and injecting noise during training naturally bias the network towards learning invariant representations.
Borel spaces, April 1988
Sterling K. Berberian · 1988
Earlier work this paper cites.
General Pattern Theory
Ulf Grenander · 1993
Earlier work this paper cites.
Keeping the neural networks simple by minimizing the description length of the weights
Geoffrey E Hinton and Drew Van Camp · 1993
Earlier work this paper cites.
Some large-scale matrix computation problems
Zhaojun Bai, Gark Fahey, and Gene Golub · 1996
Earlier work this paper cites.
Flat minima
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
The information bottleneck method
Naftali Tishby, Fernando C Pereira, and William Bialek · 1999
Earlier work this paper cites.
The elements of statistical learning , volume 1
Jerome Friedman, Trevor Hastie, and Robert Tibshirani · 2001
Earlier work this paper cites.
Learning deep architectures for ai
Yoshua Bengio · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky and Geoffrey Hinton · 2009
Earlier work this paper cites.
On the set of images modulo viewpoint and contrast changes
Ganesh Sundaramoorthi, Peter Petersen, V. S. Varadarajan, and Stefano Soatto · 2009
Earlier work this paper cites.
Classification with scattering operators
Joan Bruna and Stéphane Mallat · 2011
Earlier work this paper cites.
Elements of information theory
Thomas M Cover and Joy A Thomas · 2012
Earlier work this paper cites.
Learning invariant feature hierarchies
Yann LeCun · 2012
Earlier work this paper cites.
A pac-bayesian tutorial with a dropout bound
David McAllester · 2013
Cited alongside, same era.
Actionable information in vision
Stefano Soatto · 2013
Cited alongside, same era.
Auto-encoding variational bayes
Diederik P Kingma and Max Welling · 2014
Cited alongside, same era.
Striving for simplicity: The all convolutional net
Jost Tobias Springenberg, Alexey Dosovitskiy, Thomas Brox, and Martin Riedmiller · 2014
Cited alongside, same era.
Fast and accurate deep network learning by exponential linear units (elus)
Djork-Arné Clevert, Thomas Unterthiner, and Sepp Hochreiter · 2015
Cited alongside, same era.
Visual representations: Defining properties and deep approximations
Stefano Soatto and Alessandro Chiuso · 2016
Later among the works it cites.
Entropy-sgd: Biasing gradient descent into wide valleys
Pratik Chaudhari, Anna Choromanska, Stefano Soatto, Yann LeCun, Carlo Baldassi, Christian Borgs, Jennifer Chayes, Levent Sagun, and Riccardo Zecchina · 2017
Closest in time.
Sharp minima can generalize for deep nets
Laurent Dinh, Razvan Pascanu, Samy Bengio, and Yoshua Bengio · 2017
Closest in time.
Gintare Karolina Dziugaite and Daniel M Roy · 2017
Closest in time.
On the ability of neural nets to express distributions
Holden Lee, Rong Ge, Andrej Risteski, Tengyu Ma, and Sanjeev Arora · 2017
Closest in time.
Variational dropout sparsifies deep neural networks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Diederik P Kingma, Tim Salimans, and Max Welling · 2015
Cited alongside, same era.
Path-sgd: Path-normalized optimization in deep neural networks
Behnam Neyshabur, Ruslan R Salakhutdinov, and Nati Srebro · 2015
Cited alongside, same era.
Deep learning and the information bottleneck principle
Naftali Tishby and Noga Zaslavsky · 2015
Cited alongside, same era.
Maximally informative hierarchical representations of high-dimensional data
Greg Ver Steeg and Aram Galstyan · 2015
Cited alongside, same era.
From facial parts responses to face detection: A deep learning approach
Shuo Yang, Ping Luo, Chen-Change Loy, and Xiaoou Tang · 2015
Cited alongside, same era.
On invariance and selectivity in representation learning
Fabio Anselmi, Lorenzo Rosasco, and Tomaso Poggio · 2016
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Dmitry Molchanov, Arsenii Ashukha, and Dmitry Vetrov · 2017
Closest in time.
Structured bayesian pruning via log-normal multiplicative noise
Kirill Neklyudov, Dmitry Molchanov, Arsenii Ashukha, and Dmitry P Vetrov · 2017
Closest in time.
Exploring generalization in deep learning
Behnam Neyshabur, Srinadh Bhojanapalli, David McAllester, and Nati Srebro · 2017
Closest in time.
Opening the black box of deep neural networks via information
Ravid Shwartz-Ziv and Naftali Tishby · 2017
Closest in time.
Amortised map inference for image super-resolution
Casper Kaae Sønderby, Jose Caballero, Lucas Theis, Wenzhe Shi, and Ferenc Huszár · 2017
Closest in time.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2017
Closest in time.
Information dropout: Learning optimal representations through noisy computation
Alessandro Achille and Stefano Soatto · 2018
Closest in time.
Stochastic gradient descent performs variational inference, converges to limit cycles for deep networks
Pratik Chaudhari and Stefano Soatto · 2018
Closest in time.